← All jobs
Data Platform Administrator (Hybrid, Bangalore)
Bengaluru, India · Hybrid · Customer Success
Apply well, not just fast
Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.
About the role
KubernetesPythonBashAnsibleMachine LearningAWSGCPAzureLinuxPrometheusGrafanaNetworkingJiraConfluenceTechnical WritingCommunication
WEKA Data Platform Administrator
About the job
WEKA is architecting a new approach to the enterprise data stack built for the age of reasoning. NeuralMesh by WEKA sets the standard for agentic AI data infrastructure with a cloud and AI-native software solution that can be deployed anywhere. It transforms legacy data silos into data pipelines that dramatically increase GPU utilization and make AI model training and inference, machine learning, and other compute-intensive workloads run faster, work more efficiently, and consume less energy.
WEKA is a pre-IPO, growth-stage company on a hyper-growth trajectory. We’ve raised $375M in capital with dozens of world-class venture capital and strategic investors. We help the world’s largest and most innovative enterprises and research organizations, including 12 of the Fortune 50, achieve discoveries, insights, and business outcomes faster and more sustainably. We’re passionate about solving our customers’ most complex data challenges to accelerate intelligent innovation and business value. If you share our passion, we invite you to join us on this exciting journey.
Location: Hybrid (On-site in Bangalore 3 days; Remote 2 days)
What You’ll Be Doing
Our AI Hyperscaler customers sell GPU capacity to other companies. Their tenants run training, inference, and checkpointing on WEKA. Those tenants expect the data layer to just work.
As a WEKA Data Platform Administrator, you own that outcome. You run the WEKA deployment for these cloud companies as an extension of their team, day to day, so they do not have to staff WEKA administration itself. You may work directly with their end customers to tune the infrastructure for the workloads they actually run.
Operate the Deployment
- Run the WEKA environment day to day: monitoring, health, upgrades, capacity, configuration, and audits.
- Create and manage file systems, tiering to object storage, snapshots, quotas, and protocol configuration (POSIX, NFS, SMB, S3) as tenants onboard and grow.
- Own the operational run state. Keep the environment healthy, documented, and recoverable.
- Execute approved changes to a standard: confirm scope and approval, classify risk, validate, keep a rollback, record the outcome.
- Watch for drift, capacity pressure, and anomalies. Raise them early with a recommended path.
Optimize Infrastructure for NCP End Customers’ Workloads
- May involve working directly with the customer’s tenants and their platform teams to understand how each workload uses storage: training, inference, checkpointing, data loading, and scheduling.
- Tune WEKA to fit the workload. Filesystem layout, tiering, snapshot policy, quotas, and per-tenant QoS.
- Validate and standardize the client layer: GPU node clients, BlueField-3 and SR-IOV networking, mount options, cgroups and kubelet CPU policy, and the WEKA Kubernetes operator and CSI driver.
- Onboard new tenants and clients with a repeatable, verified procedure. Check configuration and scripts before they reach production, not during an escalation.
- Establish performance baselines. Run benchmarks and tests with tenants, and report what changed and why.
- Find scaling limits before customers do: client counts, process ceilings, NAT and routing behavior across tenant networks.
Runbooks, Automation, and Continuity
- Build and maintain customer operating runbooks that a qualified backup can follow: procedures, access guidance, validation, rollback, and recovery steps.
- Turn repeated manual work into scripts, automation, or checklists. Use Python, Bash, and Ansible.
- Prepare handovers and backup briefs before time off or coverage changes.
- Explain procedures and decisions to customer operators so they get more capable, not more dependent.
Escalation and Collaboration
- Open and update cases with the observed condition, customer impact, work completed, evidence, and requested next step. Follow the established path to Technical Services, GRE, and R&D.
- Keep the extended WEKA team informed of changes, risk, and anything the customer needs to hear.
- Work with cloud providers, network teams, and other third parties in the customer’s environment to resolve issues.
- Contribute operational results, trends, and risks to weekly status and quarterly reviews.
Requirements
- 10+ years administering enterprise infrastructure at scale, with strong storage experience (file, object, or block).
- Deep Linux administration skills, including RHEL-family and Ubuntu, in distributed environments.
- Hands-on experience with high-speed networking: Ethernet and InfiniBand, RDMA, and the fundamentals of DNS, ACLs, and TCP/IP.
- Working knowledge of Kubernetes and containers, including persistent storage, CSI, and operators.
- Experience operating GPU, HPC, or large cloud infrastructure for external customers or multiple tenants.
- Experience with at least one major cloud: AWS, Azure, Google Cloud, or Oracle Cloud.
- Proficiency in Python, Bash, or Ansible. You automate what you repeat.
- Experience with monitoring and alerting stacks (Prometheus, Grafana, log and metrics tooling).
- Experience with ticketing systems such as Jira, and discipline in major incident management.
- Clear written and verbal communication. You can explain a change to an engineer and to an executive.
- Strong technical writing. Your runbooks work for someone who has never seen the environment.
It’s Nice If You Have
- Prior experience with WEKA or another parallel or distributed file system.
- Experience running Slurm or Kubernetes-based AI/ML platforms.
- Experience with NVIDIA GPU cluster networking and DGX-class systems.
- Experience collaborating between support and product or engineering teams.
- Familiarity with Confluence, Slack, and other collaboration tools.
What Success Looks Like
- Routine work is completed predictably, documented, and verified after execution.
- A qualified backup can operate the environment from your runbook.
- Customer’s tenants get onboarded faster, with fewer surprises, and hit their performance targets.
- Client-layer problems are found in verification, not in production.
- Risks and access gaps are raised early with a recommended path forward.
- Repeated work gets safer through checklists, scripts, and automation.
The WEKA Way
We are Accountable: We take full ownership, always–even when things don’t go as planned. We lead with integrity, show up with responsibility & ownership, and hold ourselves and each other to the highest standards.
We are Brave: We question the status quo, push boundaries, and take smart risks when needed. We welcome challenges and embrace debates as opportunities for growth, turning courage into fuel for innovation.
We are Collaborative: True collaboration isn’t only about working together. It’s about lifting one another up to succeed collectively. We are team-oriented and communicate with empathy and respect. We challenge each other and conduct positive conflict resolution. We are being transparent about our goals and results. And together, we’re unstoppable.
We are Customer Centric: Our customers are at the heart of everything we do. We actively listen and prioritize the success of our customers, and every decision we make is driven by how we can better serve, support, and empower them to succeed. When our customers win, we win.