hirq
← All jobs

Nscale

Principal Software Engineer - Fleet Management

AMER · On-site · AI Infra Technology

$240K – $400K

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

PythonDistributed SystemsAirflowAWSGCPKubernetesTerraformObservabilitySREGenerative AIIncident ResponseMentoring
About Nscale Nscale is building a vertically integrated GenAI cloud from sustainable data centers to advanced AI infrastructure and enterprise applications. Our culture values open collaboration, ownership, and excellence. About the Role We're hiring a Principal Software Engineer to lead the technical development of our Fleet Manager platform - the workflow automation system that provisions, tests, and remediates GPU nodes and network switches at scale. As technical lead, you'll own the architecture and delivery of foundational Python-based automation systems that manage the entire lifecycle of our compute infrastructure: device enrolment, burn-in testing, network configuration, GPU health monitoring, and self-healing capabilities. You'll mentor a team of senior engineers, set technical direction, and drive engineering excellence while remaining hands-on with critical systems. What you'll do - Lead technical architecture and roadmap for Fleet Manager's workflow automation systems - Own end-to-end delivery of device provisioning, validation, testing, and remediation workflows at scale - Design and build workflow orchestration systems for GPU node and network switch lifecycle management - Establish engineering standards for reliability, observability, and operational excellence across all Fleet Manager services - Mentor and raise the bar for a team of senior engineers through design reviews, technical leadership, and hands-on collaboration - Drive architecture decisions balancing automation complexity, reliability, and maintainability - Integrate with infrastructure tooling: DCIMs, NetBox, OpenStack, bare metal APIs (MAAS, Ironic, IPMI) - Partner with Infrastructure, Platform, and SRE teams to translate operational needs into robust, scalable automation - Build production-grade Python systems for hardware lifecycle automation, leveraging AI tools to accelerate delivery About You - You have 12-15+ years software engineering experience building and operating production systems, with proven technical leadership in infrastructure automation or workflow tooling - Strong Python engineering fundamentals with experience leading complex, multi-service distributed systems - You are driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement - Technical expertise: quickly understanding systems design tradeoffs, keeping track of rapidly evolving software systems - Track record of owning technical roadmaps and delivering large-scale automation systems from ambiguous requirements to production - You use AI tools like Claude or Cursor as a core part of your development workflow - as a fundamental multiplier of what you can build - Deep understanding of operational excellence: SLOs, monitoring, alerting, incident response, and production reliability - Strong mentorship skills with ability to develop high-performing engineering teams - Excellent communication skills to build consensus with stakeholders, both internally and externally Strong candidates will have - Experience with workflow orchestration tools like Temporal, Airflow, Prefect, or similar - Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems - Bare metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation - Experience building hardware lifecycle automation: provisioning, validation, testing, or remediation workflows - GPU infrastructure experience: health monitoring, burn-in testing, or cluster management - HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE) - Deep knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP - Track record of 1+ years leading large-scale, complex projects or technical teams - Open-source contributions in infrastructure automation or cloud-native tooling The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $240,000—$400,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. Nscale does not accept unsolicited candidate submissions from recruitment agencies.