hirq
← All jobs

Avra

Member of Technical Staff | Observability & Reliability

São Paulo · Remote · FullTime · Engineering

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

ObservabilityKubernetesAWSGCPTerraformMachine LearningIncident Response
ABOUT THE ROLE At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. In this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments — so observability and reliability have to work the same way everywhere. WHAT YOU'LL DO - Evolve our observability stack for logs, metrics, traces, and alerting. - Make sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane. - Bring telemetry into customer clusters within a model where agents only make outbound connections. - Detect drift between the desired state and what's actually running in each environment. - Monitor the health of our deployment and runtime agents. - Provide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference. - Define SLOs, lead incident response and postmortems, and reduce MTTR — including when a fix requires coordinating with the customer. - Reduce telemetry cost: less redundant data, more useful signal. HOW WE MEASURE SUCCESS - 99.9% serving availability, with incidents trending down. - MTTR, including on-premise incidents. - Near-zero drift between desired and actual state. - All agents active and reporting, across every dataplane. WHAT WE'RE LOOKING FOR - Deep experience with OpenTelemetry and observability backends. - Hands-on practice with SLOs, error budgets, actionable alerting, and incident management. - Strong experience with Kubernetes and infrastructure as code (Terraform / Helm ). - Experience operating software in environments you don't fully control. - Production-quality code and reviews, and a willingness to operate what you build. NICE TO HAVE - Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity). - GCP or GKE, AWS or EKS. - ML multi-node/multi-cluster workloads in production. - Financial services or regulated environments.