hirq
← All jobs

selectorsoftware

Incident & Escalation Manager

Santa Clara, California · On-site · Full-Time

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

SREDistributed SystemsObservabilityMachine LearningIncident ResponseJiraProgram ManagementCommunication
About Us Selector is building an operational intelligence platform for digital infrastructure. Using an AI/ML-based analytics approach, the platform provides actionable, multi-dimensional insights to network, cloud, and application operators. It helps operations teams meet their KPIs through seamless collaboration, a search-driven conversational experience, and automated data engineering pipelines. Our solutions are used by leading Telecom, Media, Health Care, Finance, Retail, Professional Sports, and Fortune 500 enterprise organizations around the world. Our novel approach and rapidly expanding footprint position us for continued growth as a category leader. Title: Incident & Escalation Manager Location: Santa Clara HQ preferred  About the Role Selector AI is building a dedicated Incident & Escalation function to strengthen how we manage critical incidents and customer escalations as we scale. We are looking for an experienced leader who can define the vision, build the operating model, and drive execution. You will bring together existing practices across NOC, SRE, Solution Engineering, Engineering, Product, Support, and Account teams and establish a consistent, scalable approach to incident and escalation management. This is not simply an incident coordination role. You will build the function—from vision and process through execution, training, metrics, and continuous improvement What You Will Own - Define the model: Establish what is an incident vs. escalation, severity levels, entry/exit criteria, roles, decision rights, and escalation paths. - Build the process: Turn the model into practical workflows across NOC, SRE, Solution Engineering, Engineering, Product, Support, and Account teams. - Run the response: Own the pager and participate in the Incident Commander rotation, driving clear owners, actions, decisions, communication, and resolution. - Build the IC bench: Train, coach, and certify Incident Commanders so the organization does not depend on a few individuals. - Establish customer communication: Define communication cadence, executive updates, customer RCAs, and post-incident follow-through. - Drive accountability: Every escalation has an owner, next action, and date. Challenge stalled work and escalate when commitments are missed. - Close the loop: Ensure post-incident actions have owners and deadlines and are driven to completion. - Build the feedback loop: Identify recurring customer and operational issues and drive them into Engineering and Product priorities. - Measure the function: Establish meaningful metrics around response, resolution, escalations, RCA SLAs, repeat incidents, and corrective-action closure. - Evolve the model: Use automation and AI where appropriate to improve detection, triage, investigation, and incident response while maintaining clear human accountability. What You Bring - 8–12+ years in incident management, escalation management, SRE operations, technical support escalation, or technical program management. - Experience building or transforming an incident/escalation program, not simply operating within one. - Strong ability to turn an ambiguous vision into process, ownership, tooling, training, and measurable outcomes. - Experience building and developing Incident Commanders. - Strong technical understanding of distributed systems, cloud, networking, observability, or enterprise platforms. - Calm under pressure, decisive with incomplete information, and willing to challenge unclear ownership or stalled execution. - Strong written and verbal communication, including executive updates and customer-facing RCAs. - Familiarity with PagerDuty, Jira, Slack, and incident management workflows. - AIOps, NOC, networking, telecom, or AI-assisted operations experience is a strong plus. - Willingness to participate in the Incident Commander rotation, including off-hours coverage while the function is being established. What Success Looks Like You will build a predictable, scalable Incident & Escalation capability where the right people engage quickly, ownership is clear, communication is consistent, resolution is driven, and lessons become measurable improvements. This is an opportunity to build and lead an important operational capability at a growing AI company. - Compensation: 165K–200K base Perks: discretionary PTO, health insurance, 401k, bonus potential, and more.