Base Career helps you apply smarter for this job.
Key skills for this role
We are seeking a Senior Machine Learning Engineer – Synthetic Data & Document Understanding to own the synthetic data generation track within ABBYY’s Document AI Data team.
This role focuses on building generative pipelines that produce high-quality, diverse, and realistic synthetic training data at scale. You will ensure synthetic data meaningfully improves downstream model performance by maintaining strong alignment with real-world document structures, formats, and statistical properties.
This is an ideal role for engineers who combine deep generative modeling expertise with rigorous data quality evaluation and production engineering skills.
Key Responsibilities
Technical Development & Innovation
Design and implement pipelines that analyze real documents to inform high-fidelity synthetic data generation
Build generative systems capable of producing documents across diverse formats, layouts, and domains
Develop evaluation frameworks to ensure synthetic data maintains distributional fidelity and diversity
Research and apply generative modeling techniques suited for document AI training
Identify and mitigate quality issues to ensure synthetic data is effective for downstream model training
Partner with Modeling teams to measure the impact of synthetic data on model performance
Project Ownership & Leadership
Own the synthetic data generation track end-to-end, from architecture to quality validation
Drive architectural decisions balancing quality, diversity, scale, and cost efficiency
Define and maintain data quality metrics and generation dashboards
Collaborate closely with annotation teams to ensure compatibility with downstream pipelines
Contribute to roadmap planning alongside Principal-level leadership
Infrastructure & Scale
Build scalable pipelines capable of generating millions of synthetic training examples
Implement post-processing, filtering, and validation mechanisms to remove low-quality outputs
Design cost-efficient workflows balancing compute, quality, and throughput
Develop monitoring systems to detect distribution shifts or quality degradation over time
Collaborate with Platform teams on compute orchestration, storage, and scheduling
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Bengaluru, IND
London, GBR
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
London, GBR
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
, IND
We are seeking a Senior Machine Learning Engineer – Synthetic Data & Document Understanding to own the synthetic data generation track within ABBYY’s Document AI Data team.
This role focuses on building generative pipelines that produce high-quality, diverse, and realistic synthetic training data at scale. You will ensure synthetic data meaningfully improves downstream model performance by maintaining strong alignment with real-world document structures, formats, and statistical properties.
This is an ideal role for engineers who combine deep generative modeling expertise with rigorous data quality evaluation and production engineering skills.
Key Responsibilities
Technical Development & Innovation
Design and implement pipelines that analyze real documents to inform high-fidelity synthetic data generation
Build generative systems capable of producing documents across diverse formats, layouts, and domains
Develop evaluation frameworks to ensure synthetic data maintains distributional fidelity and diversity
Research and apply generative modeling techniques suited for document AI training
Identify and mitigate quality issues to ensure synthetic data is effective for downstream model training
Partner with Modeling teams to measure the impact of synthetic data on model performance
Project Ownership & Leadership
Own the synthetic data generation track end-to-end, from architecture to quality validation
Drive architectural decisions balancing quality, diversity, scale, and cost efficiency
Define and maintain data quality metrics and generation dashboards
Collaborate closely with annotation teams to ensure compatibility with downstream pipelines
Contribute to roadmap planning alongside Principal-level leadership
Infrastructure & Scale
Build scalable pipelines capable of generating millions of synthetic training examples
Implement post-processing, filtering, and validation mechanisms to remove low-quality outputs
Design cost-efficient workflows balancing compute, quality, and throughput
Develop monitoring systems to detect distribution shifts or quality degradation over time
Collaborate with Platform teams on compute orchestration, storage, and scheduling
ABBYY is an enterprise AI company founded in 1989 that develops intelligent document processing, process mining, and task mining solutions used by over 10,000 customers including many Fortune 500 companies.
Visit company websiteJobs and hiring trendsSenior · 5+ years experience
Apply faster on company sites with our extension.