MaxInsights

Machine Learning Engineer (Video Understanding Segmentation)

MaxInsights Santa Clara, CA, US
Full-time Posted 10 days ago

Responsibilities

  • check_circle Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.
  • check_circle Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.
  • check_circle Design and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.
  • check_circle Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes.
  • check_circle Architect agentic systems and orchestration pipelines that chain embedding, captioning, retrieval, and LLM reasoning steps into reliable, end-to-end video understanding workflows.
  • check_circle Develop and scale video search infrastructure (vector indexing, retrieval, ranking) to support semantic and multi-modal queries over millions of video clips.
  • check_circle Collaborate with annotation, data engineering, and robotics teams to integrate video understanding outputs into downstream training pipelines for embodied AI and robot learning.
  • check_circle Evaluate and benchmark embedding models, LLMs, and agentic frameworks against production needs; track frontier research and bring relevant techniques into the platform.
  • check_circle Contribute to internal tooling, documentation, patents, and open-source initiatives where applicable.
  • check_circle Mentor junior engineers and interns, and help shape the long-term technical roadmap for video understanding.

Basic qualifications

  • MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
  • 3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks.
  • Strong proficiency in Python and PyTorch, with solid software engineering fundamentals.
  • Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.
  • Experience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization).
  • Familiarity with agentic system design — tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops).
  • Experience working with large-scale video data pipelines and vector search/retrieval infrastructure (e.g., FAISS, Milvus, or equivalent).

Preferred qualifications

  • PhD with a research focus in video understanding, multi-modal learning, or vision-language models.
  • Experience with temporal action segmentation, action localization, or instruction-level video chunking algorithms.
  • Experience working with egocentric video datasets or head-mounted-device (HMD) captured data.
  • Track record of deploying production-scale video search or retrieval systems.
  • Experience integrating foundation or vision-language models (e.g., CLIP, VideoCLIP, RT-1/VLA variants) into perception or decision-making pipelines.
  • Publications in top-tier computer vision or ML venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, etc).
  • Experience with humanoid robotics or embodied AI data pipelines is a plus.

Tags & Focus Areas

Fulltime Machine Learning Generative Ai Ai

Ready to Apply?

Join MaxInsights and help shape the future of AI.

Save for later

About MaxInsights

Ready to Join the Team?

Apply once with DevFound — we route your profile to MaxInsights and keep you posted on matching AI roles.