39068 Sr Systems Engineer- (AWS-Cloud Watch- Mid-level)
Brilliant · United States
Job Description
Systems Engineer- AWS Cloud Watch Location: Atlanta, Ga Salary Range : $120K-$130K, plus bonus Benefits : Medical, Unlimited PTO, 401K match. Senior Systems Engineer - Observability & Resilience (AWS / CloudWatch Focus) About the Position We are seeking a highly skilled Senior Systems Engineer to join our Observability & Resilience team. This team is in the midst of a major transformation moving away from a traditional monitoring and alert-management model toward an engineering-led observability and resilience practice focused on intelligent observability, incident correlation, automated response, operational resilience, AI-assisted operations, and engineering enablement. This is not a monitoring administrator role. This engineer will help move the organization from simply receiving alerts to understanding root cause, improving resilience, reducing operational noise, and helping engineering teams build observability into their systems earlier in the software lifecycle. The ideal candidate has owned production systems, has experienced real outages, understands what it's like when critical services fail, and can translate those lessons into better observability and operational practices. Most of our environments are cloud-based, and we currently work with CloudWatch, New Relic, SolarWinds, and Splunk, relying heavily on PagerDuty and ServiceNow ITOM for ingestion. We are not seeking users of these tools we are seeking thought leaders, owners, developers, and enablers, with an emphasis on back-end capability, adoption, and user experience. Why This Role Is Different You're Helping Transform the Team The organization is deliberately moving away from a "ticket-taking" and monitoring administration model. The goal is to build standards, drive adoption, improve resilience, enable engineering teams, and create intelligent automation. Production Experience Matters We're looking for battle-hardened operators people who have carried production responsibility, been on incident bridges, supported systems during outages, and worked through high-impact production issues. You're Not Just Using Tools We're looking for builders, integrators, automators, and platform owners not simply administrators of monitoring platforms. AI Is Important but Not as a Buzzword We're not looking for someone who's built the next AI platform. We want someone who is curious about AI, experimenting with it, passionate about reducing toil through it, and interested in augmenting operational practices. Spec-Driven Operations For the first time, we can specify how the operations harness is to be used while software is being developed. Help define that standard. Core Responsibilities AWS / CloudWatch Observability (Primary Focus) Design, build, and improve CloudWatch-based monitoring, metrics, and alerting across cloud-based environments Establish and champion cloud observability patterns as the primary standard for engineering teams operating in AWS Build synthetic monitoring and alerting that scales across distributed, cloud-native systems Drive adoption of AWS monitoring best practices across engineering teams, and help define spec-driven observability standards for cloud workloads Incident Correlation & Resilience Move the organization from alert management to incident understanding Help engineering teams understand root causes, failure chains, service dependencies, and operational impact across cloud infrastructure Observability Engineering Design and improve metrics, logs, traces, alerting, and synthetic monitoring patterns Build and champion effective observability adoption patterns across engineering teams Automation & Tooling Build scripts, integrations, automation tooling, and AI-enhanced operational solutions Own tooling development within our team's GitHub repositories Automate alert responses and document monitoring solutions Engineering Relationships Partner directly with release trains, engineering managers, technical leads, and product teams Develop and own relationships with engineering teams to drive observability adoption and governance Cross-Team Education Develop a technical specialty area and help train and mentor other engineers Spread operational knowledge across the team and broader organization About You Bachelor's degree in a related discipline and 4 years' experience in a related field (or equivalent: master's + 2 years, PhD + up to 1 year, or 16 years' experience in lieu of a degree) Deep, hands-on experience with AWS CloudWatch and cloud-native monitoring/observability patterns Professional experience optimizing the integration and flow of monitoring and ITIL systems Hands-on experience with enterprise tooling such as New Relic, SolarWinds, Splunk, PagerDuty, and ServiceNow ITOM as a builder and integrator, alongside CloudWatch Professional experience writing synthetic tests in Python, Ruby, or JavaScript using Playwright, Puppeteer, or Selenium Distributed systems expertise and understanding of failure modes Deep observability experience instrumentation, metrics, logs, traces, and alerting at scale Experience building internal platforms, developer tools, or automation that scales Git/version control and CI/CD pipeline experience Infrastructure as code and API design experience Track record eliminating toil through intelligent automation Production ownership experience (on-call, incident response, observability) Systems thinking mindset understanding how components interact at scale Eager to dig into problems and bring proposed solutions to group discussion Open to feedback and able to creatively adapt multiple ideas into solutions Strong technical writing skills, including high- and low-level diagramming techniques Analytical skills and careful attention to detail Availability for rotational on-call duties outside standard business hours may be required Why This Role Is Different (Leadership & Growth) You'll be a key player transforming a team, developing key relationships with engineering teams and driving a roadmap to enable and govern solid observability and resilience patterns. You'll work with cutting-edge LLM technology to solve real production and observability problems, help define spec-driven operations standards, and grow into technical acumen while gaining exposure to leadership across all levels. Brilliant Staffing, LLC is an Equal Opportunity Employer and encourages applications from all individuals regardless of race, color, religion, gender, gender identity, sexual orientation, national origin, disability, or veteran status.
Details
| Company | Brilliant |
| Location | United States |
| Type | FULL TIME |
| Niche | general |
