hirq
← All jobs

prometheus

Senior / Staff Software Engineer, Product Infrastructure

San Francisco · Remote · FullTime · R&D

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

PrometheusPythonGoRustC++KubernetesGitDistributed Systems
About Prometheus Prometheus is building AI systems for engineering in the physical world. Behind them sits a platform that runs large, long-lived computational workloads: fleets of agents and numerical solvers, very large binary artifacts, and results that have to be correct because real hardware depends on them. The role You will build the infrastructure this platform runs on. It is a small set of hard, coupled systems, and you will own one or more of them end to end: design, implementation, operations, and on-call. - Durable orchestration. Workloads run for days across thousands of tasks. They must survive failure, checkpoint, resume, and stay steerable by humans and agents mid-flight. - Versioned artifact storage. Git-style semantics (branch, diff, merge) over large binary artifacts, on content-addressed storage with aggressive deduplication, at terabyte-to-petabyte scale. - CI and validation pipelines. Merge queues where the gates are expensive computed checks, not just unit tests, and a change is validated by actually running it. - Compute execution fabric. Scheduling heterogeneous workloads, many on GPUs, across cloud fleets: batching, prioritization, budget enforcement, and streaming results back in real time. - Agent harness. Sandboxed, reproducible, observable execution for long-running autonomous agents that write code, run tools, and spend real compute. You might be a fit if - You have 5+ years building and operating distributed systems in production; for the Staff level, 10+ years and a system you are known for. - You have seen real scale and carry the scars: at a database, streaming, or orchestration company; a multiplayer game backend; internet-scale media storage; a simulation or GPU-compute platform; or a build/CI system serving thousands of engineers. - You have gone deep in at least one of: durable workflow engines, storage engines or content-addressed storage, high-throughput schedulers, or build/merge systems. - You write production systems code (Go, Rust, or C++) and are effective in Python. - You have carried a pager for stateful systems in production. - You want your work to matter in the physical world. Strong pluses - GPU scheduling or HPC job orchestration (Slurm, Ray, Kubernetes batch). - Deterministic simulation testing or serious correctness culture (FoundationDB, TigerBeetle, Jepsen lineage). - Exposure to scientific or geometric computing. Not required; we will teach you the domain. - Experience building infrastructure for AI agents: sandboxing, session persistence, tool execution.