Base Career helps you apply smarter for this job.
Key skills for this role
We are seeking a hands-on Testing Lead to own quality and documentation for our Deep Learning, LLM, and Vision-Language Model (VLM) products. You will define how we test, measure, document, and communicate AI quality—working closely with ML, Engineering, and Product teams in a fast-paced startup environment.
This role is ideal for someone who believes clear documentation is as critical as good testing , especially for non-deterministic AI systems.
3–4 years (hands-on ownership in DL / LLM / GenAI testing)
Full-time
We are seeking a hands-on Testing Lead to own quality and documentation for our Deep Learning, LLM, and Vision-Language Model (VLM) products. You will define how we test, measure, document, and communicate AI quality—working closely with ML, Engineering, and Product teams in a fast-paced startup environment.
This role is ideal for someone who believes clear documentation is as critical as good testing , especially for non-deterministic AI systems.
Define testing strategy for LLMs, VLMs, and DL pipelines .
Create and maintain clear, lightweight documentation covering:
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
, IND
Bengaluru, IND
Bengaluru, IND
Model testing strategies and assumptions
Evaluation metrics and acceptance criteria
Known limitations, risks, and failure modes
Release readiness and quality sign-off
Ensure documentation evolves with models, data, and prompts.
Design tests for:
Prompt templates and prompt changes
RAG pipelines (retrieval quality, grounding, hallucination control)
Multi-turn conversations and long-context behaviour
Maintain golden datasets , regression test suites, and test result summaries.
Document prompt behaviour, edge cases, and known model quirks.
Test VLMs for image-text alignment, OCR, captioning, and reasoning.
Document model performance across different image types, quality levels, and domains.
Track and publish model behaviour changes between versions.
Build Python-based automation for evaluation and regression testing.
Integrate tests into CI/CD and MLOps pipelines .
Produce readable quality reports and dashboards for engineers and leadership.
Monitor and document production issues such as model/data drift and degradation .
Establish QA and documentation standards that scale with a startup.
Mentor engineers on writing testable code and meaningful documentation.
Act as the single source of truth for AI quality, testing, and known risks.
Experience with VLMs, multimodal models, or computer vision .
Exposure to RAG architectures , vector databases, and embeddings.
Familiarity with tools like LangChain, LlamaIndex, MLflow, or similar.
Experience documenting AI risks, limitations, or compliance requirements.
Private global teleradiology provider delivering remote CT, MRI, X-ray, ultrasound, and other radiology reports to hospitals.
Visit company websiteJobs and hiring trendsFull-time
Mid · 4+ years experience
Onsite
Apply faster on company sites with our extension.