hirq
← All jobs

Moonlite

Senior Software Engineer, Network Platform

Chicago, IL or Remote · Remote · Engineering

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

NetworkingLinuxPythonGoBashFastAPIgRPCKubernetesTerraformObservabilityMentoringCommunication
Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads. We provide infrastructure deployed in our facilities or co-located in yours, delivering flexible on-demand or reserved compute that feels like an extension of your existing data center. Our team of AI infrastructure specialists combines bare-metal performance with cloud-native operational simplicity, enabling research teams and enterprises to deploy demanding AI workloads with enterprise-grade reliability and compliance. Your Role: You will be fundamental to building our software-defined networking (SDN) platform that enables high-performance, isolated networking for distributed computing, model training, inference, and data-intensive workloads. Working closely with our network, infrastructure, and product teams, you'll design and implement the network orchestration and provisioning systems that manage DPU-accelerated networking, tenant isolation, and network lifecycle management, enabling researchers and engineers to access enterprise-grade networking with cloud-like simplicity. Job Responsibilities - Software-Defined Networking Architecture: Collaborate with infrastructure to design and build scalable SDN orchestration systems leveraging NVIDIA Bluefield-3 DPUs to deliver programmable, high-performance networking for AI workloads with hardware-accelerated forwarding isolation. - Network Provisioning & Lifecycle Management: Implement automated SDN provisioning systems that handle VPC creation, subnet allocation, routing configuration, and network resource lifecycle from deployment through decommissioning. - DPU Platform Engineering: Develop platform capabilities for managing Bluefield-3 DPUs including SR-IOV virtual function management, OVS offload configuration, network function deployment, and integration with compute orchestration systems. - Multi-Tenancy & Network Isolation: Build enterprise-grade network isolation using VPCs, VXLAN, and hardware-accelerated forwarding to ensure complete tenant separation while maintaining high-performance connectivity for GPU clusters and distributed workloads. - High-Performance Networking: Collaborate with infrastructure to optimize network paths for RDMA, RoCE, and GPU-to-GPU communication, ensuring minimal latency and maximum throughput for distributed training and large-scale computational workloads. - Network APIs & Integration: Develop robust APIs and SDKs for network resource management that integrate seamlessly with compute and storage platforms, enabling programmatic network provisioning and configuration. - Network Observability: Implement comprehensive network monitoring, telemetry, and troubleshooting systems that provide visibility into network performance, utilization, and tenant traffic patterns. - Security & Policy Management: Build platform network security features including security groups, firewall rules, and policy enforcement that protect tenant workloads while enabling flexible network configuration. Requirements - Experience: 5+ years in software engineering with proven experience building network platforms, SDN systems, or network automation for production environments. - Networking Expertise: Deep understanding of networking fundamentals including TCP/IP, VLANs, VXLAN, BGP, OSPF, routing protocols, and data center network architectures. - Software-Defined Networking: Background in SDN concepts, network virtualization, overlay networks, and programmable networking technologies. - Programming Skills: Experience with Shell scripting and/or Python - Linux Networking: Strong experience with Linux networking stack, including network namespaces, iptables/nftables, Open vSwitch, and kernel networking systems. - High-Performance Networking: Understanding of RDMA, RoCE, Infiniband, and low-latency networking requirements for distributed computing and GPU workloads. - Problem-Solving & Architecture: Demonstrated ability to solve complex networking performance and scalability challenges while balancing pragmatic shipping with good long-term architecture. - Autonomy & Communication: Comfortable navigating ambiguity, defining requirements collaboratively, and communicating technical decisions through clear documentation. - Commitment to Growth: Growth mindset with continuous focus on learning and professional development. Preferred Qualifications - Experience with NVIDIA Bluefield DPU programming and DOCA framework - Background with network function virtualization (NFV) and service function chaining - Knowledge of Kubernetes networking (CNI plugins, network policies, service mesh) - Experience building network control planes or SDN controllers - Familiarity with network automation frameworks and infrastructure-as-code for networking - Understanding of data center fabric architectures (spine-leaf, CLOS topologies) - Experience with network security and compliance requirements in regulated industries - Background building networking for research institutions, HPC environments, or cloud providers Key Technologies Python, Go, NVIDIA Bluefield DPUs, Open vSwitch, VXLAN, SR-IOV, RDMA, RoCE, InfiniBand, BGP, Linux networking, Terraform, FastAPI, gRPC Why Moonlite - Build Next-Generation Infrastructure: Your work will create the platform foundation that enables financial institutions to harness AI capabilities previously impossible with traditional infrastructure. - Hands-On Ownership: As an early engineer,, you'll have end-to-end ownership of projects and the autonomy to influence our product and technology direction. - Shape Industry Standards: Contribute to defining how enterprise AI infrastructure should work for the most demanding regulated environments. - Collaborate with Experts: Work alongside seasoned engineers and industry professionals passionate about high-performance computing, innovation, and problem-solving. - Start-Up Agility with Industry Impact: Enjoy the dynamic, fast-paced environment of a startup while making an immediate impact in an evolving and critical technology space. What We Offer - Competitive salary plus startup equity - On-the-job training and mentorship from experienced engineers - Opportunities for career development within a high-growth environment - Health, dental, and vision insurance, 100% covered for the employee - 401(k) with up to 6% company match - Paid time off and holidays We are an equal opportunity employer and welcome applicants from all backgrounds. #li-remote