Summary
The Junior Site Reliability Engineer helps maintain the reliability, availability, and performance of critical customer-facing systems. This onsite Seattle role focuses on proactive monitoring, incident response, root cause analysis, and improving observability and operational processes.
Responsibilities
- Monitor systems and dashboards to identify and respond to performance or availability anomalies.
- Participate in incident response, troubleshoot issues, mitigate outages, and restore service.
- Investigate incident root causes and document findings to prevent recurrence.
- Improve monitoring, logging, alerting, automation, and reliability processes.
- Support SLO and SLI tracking, CI/CD improvements, and operational documentation.
- Collaborate with software engineering and infrastructure teams to improve system resilience and scalability.
Requirements
- Bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
- Foundational knowledge of site reliability engineering, monitoring, alerting, and incident management.
- Exposure to observability and logging tools such as Prometheus, Datadog, Grafana, Splunk, or Elasticsearch.
- Basic proficiency in a programming or scripting language such as Python, Go, Bash, or Java.
- Familiarity with cloud platforms, containers, and orchestration technologies such as Docker and Kubernetes.
- Strong analytical, communication, documentation, and collaboration skills.
- Ability to work onsite in Seattle five days a week, Sunday through Thursday.