Base Career helps you apply smarter for this job.
Key skills for this role
Contribute to the design and architecture of core platform components and evaluation systems, making the load-bearing technical decisions and bearing accountability for their reliability, scalability, and long-term maintainability.
Help set the technical direction for how AI capabilities are built, evaluated, and deployed across the company, and define a coherent platform vision that scales beyond your immediate team.
Design reusable abstractions, SDKs, and services for model integration, prompt management, experimentation, and deployment that establish organization-wide patterns and reduce duplicated effort.
Help define the evaluation strategy and methodology for AI capabilities across the company – automated metrics, human-in-the-loop workflows, test set management, and benchmarking – and establish the quality standards other teams build against.
Build evaluation frameworks and developer tooling robust enough for production yet simple enough for non-specialist developers to adopt.
Establish observability standards for AI systems – quality, performance, cost, and regression signals – and build dashboards and reporting that turn those signals into actionable decisions.
Drive engineering rigor in delivery through testing discipline, reproducibility, sound experimental design, and statistically defensible measurement of model quality.
Provide technical leadership on the team's most ambiguous and highest-impact problems, scoping and sequencing work where direction is limited.
Mentor engineers and raise engineering standards through code review, design review, and leading by example.
Contribute to model and system governance practices including documentation (model cards, system cards), dataset and test-set versioning, reproducibility, and responsible-AI checks embedded directly into the platform.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Montréal, CAN
, IND
Bengaluru, IND
Bengaluru, IND
Sheffield, GBR
Mumbai, IND
Montréal, CAN
, GBR
Mumbai, IND
Act as a technical multiplier – codifying best practices into tooling and standards adopted by hundreds of developers.
Track developments in LLMs, evaluation research, and AI tooling, and translate them into pragmatic, well-scoped improvements to the platform.
Prototype and de-risk emerging techniques and tools, shepherding the promising ones from experiment to supported capability.
Champion the adoption of new platform capabilities across teams, lowering the barrier for developers to use them well.
Keep the platform current and competitive without chasing novelty for its own sake.
Partner with research, product, and localization leaders to align evaluation methodology with real-world quality and customer needs.
Influence roadmap and technical strategy beyond your immediate team, building consensus across engineering and product stakeholders.
Gather requirements from developers across the company and represent their needs in platform direction, acting as a trusted technical partner.
Communicate technical direction, trade-offs, and quality standards clearly to both technical and non-technical audiences.
Skills & Experience
Significant software engineering experience (typically 5+ years) building and operating production systems, tools, libraries, or services that many other engineers depend on, with excellent API design, reliability, and developer experience. Also, a track record with CI/CD and cloud infrastructure.
Proficiency in Python and/or another general-purpose language, with strong testing discipline
Hands-on experience building with LLMs or other ML systems (prompt engineering, fine-tuning, retrieval, model integration), with an understanding of their failure modes and tradeoffs.
Proven experience designing and leading evaluation for AI/ML systems: defining metrics and methodology, building evaluation pipelines, managing test sets, and reasoning rigorously about model quality and regressions.
Strong command of evaluation concepts — various types of metrics (accuracy, precision, recall, F1), the distinction and tradeoffs between automated and human evaluation, statistical significance, and the limits of each approach.
Excellent written and verbal communication, and a history of influencing technical direction across teams and mentoring other engineers.
Comfort with ambiguity and the judgment to scope, prioritize, and sequence high-impact work with limited direction.
Deep experience evaluating NLP, machine translation, or content-generation systems, including metrics such as COMET, chrF++, BLEU, MetricX, and MQM-style human evaluation.
Experience with experimentation and observability tooling, data/test-set versioning, and rigorous benchmarking workflows.
Established practice in AI governance and documentation - model cards, system cards, reproducibility, and responsible-AI considerations - at an organizational level.
Broad familiarity with the modern LLM ecosystem (open and proprietary models, orchestration frameworks, vector stores) and well-formed views on the tradeoffs.
Experience supporting multilingual or localization-focused products at enterprise scale.
For all applicants in Poland and accordance with applicable pay transparency legislation, candidates will be informed of the salary range for this position prior to interview. Pay is determined without reference to previous salary history and is aligned to objective job‑related criteria.
Life at RWS - If you like the idea of working with smart people who are passionate about growing the value of ideas, data and content by making sure organizations are understood , then you’ll love life at RWS.
Our purpose is to unlock global understanding. This means our work fundamentally recognizes the value of every language and culture. So, we celebrate difference, we are inclusive and believe that diversity makes us strong. We want every employee to grow as an individual and excel in their career.
In return, we expect all our people to live by the values that unite us: to partner , putting clients fist and winning together , to pioneer , innovating fearlessly and leading with vision and courage, to progress , aiming high and growing through actions and to deliver , owning the outcome and building trust with our colleagues and clients.
RWS embraces DEI and promotes equal opportunity, we are an Equal Opportunity Employer and prohibit discrimination and harassment of any kind. RWS is committed to the principle of equal employment opportunity for all employees and to providing employees with a work environment free of discrimination and harassment. All employment decisions at RWS are based on business needs, job requirements and individual qualifications, without regard to race, religion, nationality, ethnicity, sex, age, disability, or sexual orientation. RWS will not tolerate discrimination based on any of these characteristics.
Get the 3Ps right – Partner, Pioneer, Progress – and we´ll Deliver together as RWS.
Recruitment Agencies: RWS Holdings PLC does not accept agency resumes. Please do not forward any unsolicited resumes to any RWS employees. Any unsolicited resume received will be treated as the property of RWS and Terms & Conditions associated with the use of such resume will be considered null and void.
Public UK AI solutions company providing enterprise localization, content management, linguistic data, and intellectual-property services.
Visit company websiteJobs and hiring trendsFull-time
Senior · 5+ years experience
Remote
Apply faster on company sites with our extension.