hirq
← All jobs

Lwolf

Cloud Operations Team Lead

Manila, Philippines · Hybrid · Full-time · IT

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

AzureAWSSRETerraformAnsibleIncident ResponseAgile
JOB SUMMARY: Be part of a fast moving, lean product development Organization, supporting the Lone Wolf family of Applications. Serve as the hands-on Team Lead for a Cloud Operations team responsible for the reliability, performance, changes, releases, maintenance, and day-to-day operation of infrastructure across AWS, Azure, and a private data center. Guide and mentor a small team of engineers while remaining actively engaged in the work, partnering with Development, Product, IT, DBA, and SRE teams to keep Lone Wolf's hybrid, multi-cloud environment stable, secure, and scalable across all product lines. ESSENTIAL RESPONSIBILITIES AND DUTIES: - Lead a team of Cloud Operations Engineers, providing day-to-day technical direction, prioritization, and hands-on support - Mentor and develop engineers by sharing best operational practices and deepening the team's understanding of the applications and infrastructure they support - Help build a best-in-class team in a hybrid and multi-cloud environment spanning AWS, Azure, and a private data center - Operate, monitor, and maintain infrastructure and platform resources across AWS, Azure, and on-premises data center systems - Recommend and help implement tooling that drives operational efficiencies - Support orchestration, self-service, infrastructure-as-code, and messaging platforms in partnership with the SRE and DBA teams - Manage and conduct infrastructure changes and application releases as per Release Management team schedule and needs - Manage and conduct proactive maintenance of systems to ensure compliance and stability - Respond to and help lead resolution of operational incidents, coordinating across teams and vendors as needed - Provide analysis of operational efficiencies gained from increased automation, and evaluate new tools and processes to further improve operational and development outcomes. ​ OFF-HOURS / ON-CALL EXPECTATIONS: - Participate in a rotating on-call schedule providing after-hours and weekend coverage for Cloud Operations incidents and escalations - Serve as an escalation point during off-hours for incidents that require hands-on infrastructure expertise across AWS, Azure, or the private data center - Coordinate off-hours maintenance windows, planned releases, and change activity in partnership with the Release Management team - Help establish and refine on-call runbooks, escalation paths, and handoff procedures to reduce mean time to resolution. QUALIFICATIONS: - 5+ years of operating and maintaining infrastructure and platforms in AWS, Azure, and private data centers - 5+ years of building and maintaining automation - 2+ years of proven leadership experience - Experience with AI driven tools to assist with analysis, changes, and troubleshooting - Experience with tools such as K8, Ansible, Terraform - Experience with Azure Service Bus and queueing technologies - In-depth understanding of network concepts - In-depth understanding of security practices - Background in an infrastructure role: Server, Network, Storage, etc. - Experience working for a high growth SaaS company - Agile, lean, or similar development methodologies/experience - Experience carrying a production on-call rotation, including after-hours incident response and escalation KNOWLEDGE/SKILLS/ABILITIES: - Ability to use a wide variety of open-source technologies and tools - Ability to leverage AI platforms to enhance individual/team performance - Ability to code and script - Experience with systems and IT operations - Comfort with frequent, incremental code testing and deployment - Strong grasp of automation tools - Data management skills - A strong focus on business outcomes - Comfort with collaboration, open communication and reaching across functional borders - Process re-engineering - Communicating and collaborating with others - Experience across multiple cloud providers (AWS and Azure) a plus - Self-motivated and able to work in a small team with minimal supervision - Desire to deliver the best solutions for the need - Ownership of any output delivered - Ability to justify decisions made within the work being delivered - Ability to use data to drive decisions