← All jobs
Senior DevOps Engineer
Singapore, Central Singapore, Singapore · Remote · fulltime_permanent · Developers
Apply well, not just fast
Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.
About the role
CI/CDKubernetesObservabilitySREIncident ResponseGCPIAMAuditing
Senior DevOps Engineer
Location: Singapore or Kuala Lumpur
Employment Type: Full-time
About the Opportunity
Nodeworthy is supporting a leading Solana-based decentralized exchange in hiring a Senior DevOps Engineer.
You will own and improve the infrastructure, deployment systems and operational reliability behind a high-traffic production platform. The role covers cloud infrastructure, Kubernetes, CI/CD, observability, incident response, infrastructure automation and production security.
This is a deeply hands-on individual-contributor position. You should be comfortable designing infrastructure, writing automation, deploying changes, investigating production incidents and taking ownership of systems from implementation through ongoing operation.
The ideal candidate is an experienced DevOps or SRE engineer who also possesses strong security fundamentals. You should be able to keep production systems reliable and scalable while ensuring that infrastructure, access controls, deployment processes and developer environments remain appropriately secured.
This role does not cover smart-contract auditing or protocol-level security reviews, which are handled by a separate internal team.
What You’ll Own
Cloud Infrastructure and Platform Engineering
- Design, build and operate highly available production infrastructure on Google Cloud Platform.
- Manage and improve Kubernetes clusters, networking, compute, storage, databases and other managed cloud services.
- Build reusable infrastructure through infrastructure-as-code and automation.
- Improve the scalability, reliability, performance and cost efficiency of production systems.
- Establish consistent infrastructure standards across development, staging and production environments.
- Identify and resolve infrastructure bottlenecks, configuration issues and operational risks.
- Evaluate and implement new infrastructure tools where they provide clear operational value.
- Maintain clear documentation and runbooks for critical systems.
CI/CD and Deployment Automation
- Own and improve CI/CD pipelines, build systems and deployment workflows.
- Automate testing, builds, releases, environment provisioning and production deployments.
- Improve deployment speed and reliability while maintaining appropriate safeguards.
- Secure source-code repositories, build environments, automation systems and deployment credentials.
- Implement branch protections, approval workflows, rollback processes and deployment auditability.
- Review infrastructure-as-code and deployment configurations for reliability and security issues.
- Work closely with engineers to make deployments simple, repeatable and safe.
Reliability, Monitoring and Observability
- Build and maintain monitoring, logging, alerting and observability across infrastructure and applications.
- Define service-level indicators and operational metrics for critical production systems.
- Improve visibility into system health, application performance and infrastructure utilisation.
- Detect and investigate abnormal service behaviour before it becomes a major incident.
- Reduce alert noise and ensure that operational alerts are clear and actionable.
- Identify recurring production issues and implement long-term preventive improvements.
- Conduct capacity planning and prepare systems for periods of increased activity.
Production Operations and Incident Response
- Respond to production incidents involving infrastructure, services, APIs, databases and deployment systems.
- Lead technical investigations from initial alert through containment, diagnosis and resolution.
- Coordinate recovery efforts across DevOps, SRE, backend and frontend engineering.
- Perform root-cause analysis and lead post-incident reviews.
- Develop and maintain incident-response and disaster-recovery playbooks.
- Improve system resilience, redundancy, backup and recovery processes.
- Participate in an appropriate production support or on-call rotation.
Infrastructure and Operational Security
- Apply security best practices across cloud infrastructure, Kubernetes, CI/CD and developer environments.
- Manage IAM, service accounts, privileged access and least-privilege policies.
- Improve the management of credentials, secrets, API keys and deployment permissions.
- Monitor infrastructure and production systems for suspicious access, configuration changes and abnormal activity.
- Identify exposed services, excessive permissions, insecure configurations and other infrastructure weaknesses.
- Introduce network segmentation and appropriate controls around sensitive systems.
- Strengthen software-supply-chain security across repositories, dependencies and build environments.
- Support the investigation and remediation of infrastructure or application security incidents.
- Establish secure infrastructure baselines and monitor environments for deviations.
Engineering Collaboration
- Partner with backend and frontend engineers to improve application deployment, reliability and performance.
- Provide infrastructure guidance during the design of new services and products.
- Help engineers diagnose production and performance issues.
- Review backend architecture from an operational, scalability and security perspective.
- Contribute to backend or DevOps code when additional engineering support is required.
- Improve service-to-service authentication, API security and production configuration.
- Make infrastructure and security practices practical and easy for engineers to adopt.