Machine Learning Engineer (Video Understanding Segmentation)
Actively Hiring
Full-time Posted 10 days ago
Responsibilities
- check_circle Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.
- check_circle Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.
- check_circle Design and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.
- check_circle Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes.
- check_circle Architect agentic systems and orchestration pipelines that chain embedding, captioning, retrieval, and LLM reasoning steps into reliable, end-to-end video understanding workflows.
- check_circle Develop and scale video search infrastructure (vector indexing, retrieval, ranking) to support semantic and multi-modal queries over millions of video clips.
- check_circle Collaborate with annotation, data engineering, and robotics teams to integrate video understanding outputs into downstream training pipelines for embodied AI and robot learning.
- check_circle Evaluate and benchmark embedding models, LLMs, and agentic frameworks against production needs; track frontier research and bring relevant techniques into the platform.
- check_circle Contribute to internal tooling, documentation, patents, and open-source initiatives where applicable.
- check_circle Mentor junior engineers and interns, and help shape the long-term technical roadmap for video understanding.
Basic qualifications
- MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
- 3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks.
- Strong proficiency in Python and PyTorch, with solid software engineering fundamentals.
- Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.
- Experience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization).
- Familiarity with agentic system design — tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops).
- Experience working with large-scale video data pipelines and vector search/retrieval infrastructure (e.g., FAISS, Milvus, or equivalent).
Preferred qualifications
- PhD with a research focus in video understanding, multi-modal learning, or vision-language models.
- Experience with temporal action segmentation, action localization, or instruction-level video chunking algorithms.
- Experience working with egocentric video datasets or head-mounted-device (HMD) captured data.
- Track record of deploying production-scale video search or retrieval systems.
- Experience integrating foundation or vision-language models (e.g., CLIP, VideoCLIP, RT-1/VLA variants) into perception or decision-making pipelines.
- Publications in top-tier computer vision or ML venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, etc).
- Experience with humanoid robotics or embodied AI data pipelines is a plus.
Tags & Focus Areas
Fulltime Machine Learning Generative Ai Ai
About MaxInsights
Ready to Join the Team?
Apply once with DevFound — we route your profile to MaxInsights and keep you posted on matching AI roles.