MLOps / LLMOps Engineer (GenAI Platform) $145k - $156k in Santa Clara at Unitedone health
U

MLOps / LLMOps Engineer (GenAI Platform)

Unitedone health Santa Clara, CA, US
Full-time $145k - $156k Posted 2 days ago

Role overview

We are seeking an experienced MLOps / LLMOps Engineer to design, deploy, and optimize production-grade Generative AI and Large Language Model (LLM) platforms. The ideal candidate will have strong expertise in Python, AI/ML platform engineering, model serving, Kubernetes, and cloud-native MLOps practices.

Responsibilities

  • check_circle Design, build, and maintain scalable AI/ML infrastructure for enterprise LLM applications.
  • check_circle Deploy and optimize LLM inference workloads for high performance and low latency.
  • check_circle Implement scalable model serving, monitoring, and observability solutions.
  • check_circle Collaborate with AI researchers, data scientists, and software engineers to deliver production-ready GenAI solutions.
  • check_circle Improve GPU utilization, model performance, and operational efficiency.
  • check_circle Ensure AI governance, security, and Responsible AI compliance across deployments.

Basic qualifications

  • 5–7 years of experience in <strong>MLOps, LLMOps, AI/ML Platform Engineering, or Machine Learning Engineering</strong>.
  • Strong proficiency in <strong>Python</strong> and software engineering best practices.
  • Hands-on experience with <strong>open-source LLMs</strong> such as <strong>Llama, Mistral, Gemma, or Qwen</strong>.
  • Expertise in <strong>LLM inference and model hosting</strong> using technologies such as:
  • vLLM
  • SGLang
  • Hugging Face TGI
  • NVIDIA Triton Inference Server
  • Ray Serve
  • Azure Machine Learning
  • Databricks Model Serving
  • Experience with <strong>Kubernetes, Docker, Azure ML, Databricks, and MLflow</strong>.
  • Strong understanding of:
  • Retrieval-Augmented Generation (RAG)
  • Vector Databases
  • GPU Optimization
  • Model Quantization
  • KV Cache
  • PagedAttention
  • Continuous/Dynamic Batching
  • Proven experience building, deploying, troubleshooting, scaling, and optimizing <strong>production-grade GenAI and LLM applications</strong>.
  • Experience implementing AI observability, governance, and Responsible AI best practices.

Preferred qualifications

  • Hands-on experience with <strong>LLM Fine-Tuning</strong> using:
  • PEFT
  • SFT
  • CPT
  • LoRA
  • QLoRA
  • Experience with:
  • Azure AI Foundry
  • Azure OpenAI
  • Hugging Face
  • DeepSpeed
  • PEFT
  • Knowledge of distributed training and multi-GPU environments.
  • Experience with Agentic AI frameworks such as:
  • LangGraph
  • AutoGen
  • CrewAI
  • Familiarity with simulation platforms, digital twins, scientific computing, or modeling and simulation workflows.

Benefits

  • check_circle Work on cutting-edge Generative AI and Large Language Model technologies.
  • check_circle Build enterprise-scale AI platforms using modern cloud-native tools.
  • check_circle Collaborate with highly skilled AI and ML engineering teams on innovative projects.

Tags & Focus Areas

Fulltime Ai Machine Learning Mlops Generative Ai

Ready to Apply?

Join Unitedone health and help shape the future of AI.

Save for later

About Unitedone health

Ready to Join the Team?

Apply once with DevFound — we route your profile to Unitedone health and keep you posted on matching AI roles.