Site Reliability Engineer
Nscale · United States
Job Description
About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each other the truth, and everyone here stays close to the infrastructure that makes AI work. The Role This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll be expected to make the systems you touch quieter over time. What You'll Do • Build and own the automation and tooling that keeps the platform running; treat operational toil as a bug to be fixed, not a fact of life. • Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance. • Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post- incident reviews that actually change the system. • Investigate performance and reliability problems across Linux, networking, and distributed services, then fix them at the source. • Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the stack. • Improve availability, scalability, and efficiency through code, not manual effort. What You'll Bring • 3-6 years in SRE, systems engineering, or software engineering, including time running production in a data center or cloud environment. • Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work away.
- Solid command of Linux, networking fundamentals, and distributed systems.
- A track record of troubleshooting live production issues and owning the fix
- Experience with AI or GPU workloads, or high-performance computing (HPC).
- Familiarity with high-performance networking (InfiniBand, RDMA).
- Kubernetes, plus virtualized or bare-metal environments.
- Competitive base plus equity, reviewed every 12 months.
- Real scope early, and a progression plan built around the skills you want to
$100,000—$170,000 USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. [Details
| Company | Nscale |
| Location | United States |
| Type | FULL TIME |
| Niche | tech |
| Experience | full-time |
