AI Engineer
Responsibilities
- check_circle Design and build high-performance inference serving systems for large-scale transformer and multimodal models (including 100B+ and MoE architectures)
- check_circle Implement and tune inference optimisations: speculative decoding, continuous batching, KV cache management, prefill/decode disaggregation, and quantisation (INT4/INT8/FP8)
- check_circle Contribute to and customise inference frameworks (vLLM, TensorRT-LLM, SGLang, or equivalent) for Zoom's production requirements
- check_circle Write and profile CUDA kernels and custom ops where framework-level optimisation is insufficient
- check_circle Own end-to-end deployment: from model packaging and serving API design to latency SLO monitoring and incident response
- check_circle Partner with research to translate model architecture changes into inference-efficient implementations
- check_circle Drive technical design and set the bar for inference engineering practices across the team
- check_circle A Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
- check_circle 5+ years of software engineering experience, with significant time spent on inference systems or ML infrastructure at production depth
- check_circle Hands-on experience with at least one major inference framework: vLLM, TensorRT-LLM, SGLang, or ONNX Runtime (serving, not just export)
- check_circle GPU programming experience: CUDA kernel development, memory optimisation, and profiling with Nsight or equivalent tools
- check_circle Production experience serving LLMs or large vision models - you've owned latency SLOs, debugged throughput regressions, and shipped optimisations that moved the needle
- check_circle Depth in at least two of: speculative decoding, continuous batching, KV cache design, quantisation pipelines, prefill/decode disaggregation
- check_circle Strong systems instincts in Python and C++; ability to read and modify framework internals
Preferred qualifications
- Advanced degree (Master's or PhD) in a relevant technical field
- Experience with MoE models or 100B+ parameter deployments
- Familiarity with disaggregated serving architectures or multi-node inference
- Background in compiler-level optimisation (XLA, Triton, or similar)
About the company
You will join a dynamic AI Infrastructure team focused on enabling high-performance AI across Zoom's products and services. The team builds the core systems that support model training, deployment, and inference at scale, driving innovation in areas such as real-time communication, computer vision, and natural language understanding.
Tags & Focus Areas
About Zoom Communications
We're seeking an experienced Machine Learning Engineer specializing in LLMs and autonomous agents and agentic AI to join our AI team. You'll be instrumental in developing and deploying intelligent agent systems that can understand context, make decisions, and execute complex tasks across our platform.
Ready to Join the Team?
Apply once with DevFound — we route your profile to Zoom Communications and keep you posted on matching AI roles.