{bc}
indeed

Senior LLM / Generative AI Engineer

INFRA ASSURE
Dubai, UAE
Full-time
Senior
Onsite
AED 12,000/month
Discovered 4 weeks ago
Large Language Models and Generative AIPythonRAG, embeddings, semantic search, and vector databasesLLM fine-tuning with LoRA, QLoRA, and PEFTLLM inference and serving with vLLM, TensorRT-LLM, TGI, or equivalentLangChain, LangGraph, LlamaIndex, Semantic Kernel, or similar frameworks
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Large Language Models and Generative AIPythonRAG, embeddings, semantic search, and vector databases
Smart Apply

Full Job Posting

Role Overview

INFRA ASSURE is hiring a Senior LLM / Generative AI Engineer to own end-to-end engineering of production-grade LLM solutions.

The role spans model selection, prompt engineering, RAG, fine-tuning, AI agents, deployment, LLMOps, evaluation, security, optimization, and production operations.

The engineer will work across AI engineering, cloud and GPU infrastructure, DevOps, security, and system architecture.

Key Responsibilities

  • Evaluate and select commercial and open-source LLMs against accuracy, latency, context length, cost, licensing, data residency, and governance requirements.
  • Design prompt engineering strategies, few-shot prompts, structured outputs, function or tool calling, and context management.
  • Build RAG pipelines, semantic search, document understanding, summarization, embeddings, chunking, metadata extraction, and vector search.
  • Implement LoRA, QLoRA, PEFT, and fine-tuning for LLM customization.
  • Build agentic and multi-agent workflows with tool or API integration, memory, state management, and human approvals.
  • Deploy and operate LLM inference platforms across cloud, hybrid, and GPU environments using technologies such as vLLM, TensorRT-LLM, or TGI.
  • Optimize inference with quantization, batching, KV-cache management, speculative decoding, and prompt caching.
  • Implement LLMOps practices including versioning, deployment pipelines, rollback, canary releases, and A/B testing.
  • Develop LLM evaluation, regression testing, hallucination and groundedness checks, and LLM-as-a-judge workflows.
  • Implement observability for performance, usage, failures, refusals, quality regression, and model drift.
  • Optimize infrastructure and costs through routing, caching, right-sizing, and capacity planning.
  • Implement AI security and responsible AI controls, mentor engineers, contribute to architecture reviews, and collaborate with cross-functional teams.

Required Skills and Experience

  • 7–12+ years of experience in software engineering, AI/ML, or a related field.
  • Strong hands-on experience with LLMs, Generative AI, and production LLM-powered enterprise applications.
  • Strong Python programming and API or microservices development experience.
  • Experience with RAG, embeddings, vector databases, semantic search, and LLM applications.
  • Experience with LangChain, LangGraph, LlamaIndex, Semantic Kernel, or a similar framework.
  • Experience with LLM fine-tuning and LoRA, QLoRA, or PEFT.
  • Experience with LLM inference and serving technologies such as vLLM, TensorRT-LLM, TGI, or an equivalent.
  • Strong understanding of Docker, Kubernetes, CI/CD, Git, cloud environments, and GPU-based infrastructure.
  • Experience with AWS, Azure, or GCP.
  • Understanding of LLM evaluation, monitoring, observability, and LLMOps or MLOps.
  • Strong understanding of AI security risks including prompt injection, data leakage, and insecure tool calling.
  • Excellent problem-solving, architecture, communication, and technical leadership skills.

Preferred Experience

  • Experience with open-source models such as Llama, Mistral, Qwen, or an equivalent.
  • Experience with MLflow, Weights & Biases, LangSmith, RAGAS, DeepEval, promptfoo, or similar tools.
  • Experience with GPU optimization and inference performance tuning.
  • Knowledge of the OWASP Top 10 for LLM Applications.
  • Experience with enterprise AI governance, privacy, and responsible AI frameworks.

Candidate Profile

The preferred candidate can take an LLM solution from model selection through prompt and RAG development, fine-tuning, evaluation, agent or API integration, deployment, monitoring, optimization, and production operations.

Strong hands-on production LLM engineering experience is preferred over experience limited to basic ChatGPT or API integrations or proof-of-concept chatbots.

Compensation

  • Pay is AED 12,000–AED 15,000 per month.

Workplace

  • Work location is in person in Dubai.
  • Dubai location preference and willingness to travel 100% are listed as preferred.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at INFRA ASSURE