Base Career helps you apply smarter for this job.
Key skills for this role
You will be the language architect of Vibecoderz. The TutorAgent and its family of sub-agents only think as clearly as the prompts that guide them — and you will design those “thought patterns.” As the founding Prompt Engineer, you’ll build a prompt library and evaluation system that ensures Vibecoderz produces consistent, safe, and high-quality artifacts: slides, quizzes, code snippets, and runnable mini-apps.
This is a role for a builder of mental scaffolding. You’ll be deeply involved in orchestrating system prompts, tool schemas, chaining strategies, and eval frameworks that make our AI both reliable and delightful. You’ll collaborate with AI Engineers on orchestration, Backend Engineers on schema integration, and PM on defining product outcomes.
You’ll use Linear for execution, Notion for prompt specs and experiments, and GitHub for version control , making every iteration observable, testable, and reproducible.
You will be the language architect of Vibecoderz. The TutorAgent and its family of sub-agents only think as clearly as the prompts that guide them — and you will design those “thought patterns.” As the founding Prompt Engineer, you’ll build a prompt library and evaluation system that ensures Vibecoderz produces consistent, safe, and high-quality artifacts: slides, quizzes, code snippets, and runnable mini-apps.
This is a role for a builder of mental scaffolding. You’ll be deeply involved in orchestrating system prompts, tool schemas, chaining strategies, and eval frameworks that make our AI both reliable and delightful. You’ll collaborate with AI Engineers on orchestration, Backend Engineers on schema integration, and PM on defining product outcomes.
You’ll use Linear for execution, Notion for prompt specs and experiments, and GitHub for version control , making every iteration observable, testable, and reproducible.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Hyderabad, IND
Hyderabad, IND
Hyderabad, IND
Hyderabad, IND
Hyderabad, IND
Hyderabad, IND
Hyderabad, IND
Hyderabad, IND
Deliver prompt library v1 for TutorAgent, PlannerAgent, and QuizAgent.
Build baseline LangSmith eval pipeline with at least 20 golden prompts.
Achieve <15% hallucination rate in core flows.
Maintain <5% hallucination rate across all core workflows.
Prompt system powering 10+ specialized agents in production.
Automated regression evals integrated into CI/CD pipeline.
AI outputs powering >50% of artifacts with consistent schema compliance.
Contributions to LangChain, CrewAI, or LangSmith open source.
Research background in prompt optimization, LLM evals, or safety.
Prior work on developer-facing AI tutor or EdTech products.
Startup/founding engineer experience.
Prompting & Orchestration: LangSmith, Langfuse, Google ADK
Models: Gemini Pro, Gemini Flash, Gemini Vision, Gemini Live API, Claude 3 Opus
Data: Firestore, Redis, Neo4j (schema enforcement targets)
Infra: Cloud Run, Pub/Sub for agent comms
CI/CD: GitHub Actions with prompt regression evals
Tools: Linear (execution), Notion (prompt specs), GitHub (prompt versions)
Objective: Validate ability to design, evaluate, and optimize a prompt system powering multi-agent tutoring.
Design system prompts for: TutorAgent: teaching “React Hooks.” QuizAgent: generating 5 MCQs with answers + explanations. CodeAgent: generating runnable JS snippets.
TutorAgent: teaching “React Hooks.”
QuizAgent: generating 5 MCQs with answers + explanations.
CodeAgent: generating runnable JS snippets.
Implement prompt chaining strategy for: Outline → Lesson → Quiz → Mini-App.
Outline → Lesson → Quiz → Mini-App.
Build a LangSmith eval pipeline : Test against 20 golden prompts. Measure accuracy, latency, hallucination rate, schema compliance.
Test against 20 golden prompts.
Measure accuracy, latency, hallucination rate, schema compliance.
Optimize: Compare Gemini Flash vs. Pro routing for cost/latency. Document tradeoffs and improvements.
Compare Gemini Flash vs. Pro routing for cost/latency.
Document tradeoffs and improvements.
Prompt library (YAML/JSON).
Eval report with metrics table (baseline vs. optimized).
GitHub repo with eval scripts + LangSmith integration.
Notion doc summarizing prompt strategies.
5-min Loom walkthrough of the workflow.
Prompt Design & Schema Compliance (30%)
Prompt Chaining & Orchestration Strategy (20%)
Evaluation Pipeline & Metrics (20%)
Performance Optimization (15%)
Documentation & Reproducibility (15%)
Verified company details for this employer are not available yet.
USD 24000-32000 yearly / year
Full-time
Senior · 10+ years experience
Remote
Apply faster on company sites with our extension.