AI/ML LLM Engineer
The mission
inCall is an AI-powered receptionist for NHS GP practices. Patients call in, the AI verifies their identity against NHS records, understands their reason for calling, and handles the conversation start to finish. You are building the large language model that powers that experience.
This engagement has a two-track structure. From day one, OpenAI GPT is integrated into the live call pipeline so the full voice flow is testable immediately. In parallel, you build a fine-tuned LLaMA 3 8B model from scratch — architecture, 10,000+ synthetic NHS GP training examples, LoRA/SFT fine-tuning, RLHF/DPO, and SageMaker deployment. By the end of Month 3, the LLaMA model replaces OpenAI in the live pipeline.
What you'll build
Month 1 - OpenAI GPT integrated into the 3CX call pipeline. AWS Transcribe (en-GB) with custom NHS GP vocabulary connected for real-time STT. ElevenLabs integrated for TTS. LLaMA 3 architecture documented. Training data pipeline initiated with 10,000+ synthetic examples.
Month 2 - LLaMA 3 8B (or Mistral 7B) downloaded from Hugging Face and registered in SageMaker Model Registry. LoRA fine-tuning completed on synthetic NHS GP dataset using Hugging Face Transformers and PEFT on ml.g4dn.xlarge. SageMaker endpoint incall-poc-llama-endpoint deployed achieving minimum 70% accuracy on 100-example test set. Human evaluation of GPT vs LLaMA conversational quality conducted.
Month 3 - NHS PDS FHIR API integrated for patient context. All eight core NHS GP call scenarios demonstrated using the LLaMA model. OpenAI GPT replaced by LLaMA SageMaker endpoint in the live pipeline. Full production readiness report and technical documentation delivered.
What you bring
Hands-on experience building, training, or fine-tuning transformer-based LLMs (LLaMA, Mistral, or similar). Practical experience with LoRA, SFT, and RLHF/DPO using Hugging Face Transformers and PEFT. Strong Python and experience building synthetic data generation pipelines. AWS SageMaker for model training, hosting, and inference. Exposure to real-time STT/TTS pipelines. Comfortable working independently with daily standups and Jira-based task tracking.
The bigger picture
This is Stage 1 of a four-stage inCall product. Stage 2 adds AI triage and auto-routing. Stage 3 adds clinical routing support. Stage 4 is a fully autonomous AI call handler. Commercial model is cost-per-call licensing across GP practices. An AI engineer who delivers a clean, well-documented LLaMA POC is well-placed for the next contract.
Don't apply if
You have only used OpenAI APIs and have never built or fine-tuned a base model. You are currently engaged in other contracts that will compete for your time. You cannot commit to the daily standup and monthly deliverable rhythm.
How we hire
We receive far more applications than we can interview.
- Application form - personal information, professional screening, and five written responses answered thoroughly.
- Video submission - if selected you will receive an email within 5 working days with specific requirements. 48 hours to submit.
- Group interview - live Google Meet session with other shortlisted candidates.
- One-to-one - deep technical interview covering LLM architecture, fine-tuning approach, SageMaker deployment, and your specific plan for the inCall LLaMA build.
- Offer and contract - full contract with Key Contract Information summary included.
To apply: Submit your CV above, then complete our full application form at:
https://ladybirdltd.jp.larksuite.com/share/base/form/shrjpwYz1RbzWxvm4rVRRlkYngf
Your Indeed application alone is not sufficient - applications without the completed form will not progress.
Pay: £700.00-£3,000.00 per month
Work Location: Remote
Tags & Focus Areas
About https://ladybirdltd.com
Ready to Join the Team?
Apply once with DevFound — we route your profile to https://ladybirdltd.com and keep you posted on matching AI roles.