{bc}
linkedin

LLM / GenAI Engineer

Evlo AI
Chicago, USA
Full-time
Mid-Senior
Onsite
Discovered 1 weeks ago
PythonRAG architecturesLangChain or LlamaIndexVector databasesLLM evaluationPrompt engineering
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

PythonRAG architecturesLangChain or LlamaIndex
Smart Apply

Full Job Posting

Role Overview

The LLM / GenAI Engineer builds production AI systems combining foundation models, retrieval, structured data, and reliable software services.

The role moves GenAI systems from prototype to production while balancing answer quality, latency, cost, security, observability, and failure recovery.

Responsibilities

  • Design and deploy RAG pipelines using Python, LangChain, LlamaIndex, or custom orchestration frameworks.
  • Build secure agentic workflows connecting LLMs to APIs, databases, search systems, and business tools.
  • Develop evaluation systems using benchmark data, golden responses, LLM judges, human review, and regression testing.
  • Fine-tune and optimize models using supervised fine-tuning, LoRA or QLoRA, prompt optimization, distillation, and model routing.
  • Implement production services with Python, FastAPI, Docker, and cloud infrastructure.
  • Integrate commercial and open-source model providers.
  • Monitor token usage, model quality, hallucinations, latency, failures, and data drift with tracing, metrics, and alerts.
  • Establish controls for PII, prompt injection, data access, model versioning, and safe deployment rollbacks.

Required Qualifications

  • Three to eight years of software engineering, machine learning engineering, or applied AI experience, including at least one year delivering LLM or GenAI systems to production.
  • Strong Python and backend service experience, including asynchronous workflows, REST APIs, automated tests, and CI/CD.
  • Hands-on experience with RAG, embedding models, vector databases such as Pinecone, Weaviate, Milvus, Chroma, or pgvector, and hybrid retrieval.
  • Knowledge of LLM evaluation, prompt engineering, function calling, structured generation, fine-tuning, and quality, latency, and cost tradeoffs.
  • Experience with AWS, GCP, or Azure and containerized deployment using Docker and Kubernetes or an equivalent platform.
  • Bachelor's or master's degree in a relevant technical field, or equivalent professional experience.

Preferred Qualifications

  • Experience with Llama or Mistral, inference optimization, distributed training, multimodal models, graph-based retrieval, ML observability, responsible AI, or security practices.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today