lever
Data Engineer
Ifm-us
Sunnyvale, USA
Full-time
Mid
Onsite
USD 150000-450000 yearly / year
Discovered 1 weeks ago
PythonWeb CrawlingLarge Language ModelsCloudSparkKafka
Free
Job Fit Check
Base Career helps you apply smarter for this job.
?%
Ready to ScanKey skills for this role
PythonWeb CrawlingLarge Language Models
Key Skills for This Role
PythonWeb CrawlingLarge Language ModelsCloudSparkKafka
Full Job Posting
- Rapidly collect, curate, and preprocess datasets based on detailed specifications provided by NLPresearchers,delivering data within tight timelines.
- Develop and maintain efficient web crawling solutions, APIs, and automated workflows to continuously improve data collection processes.
- Refine and evaluate outputs from Large Language Models (LLMs) to generate structured datasets suitable for model training and benchmarking.
- Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams.
- Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria.
- Document data collection methodologies, dataset characteristics, and pipeline architecture clearly and effectively.
- Engage with peer teams and participate in technical reviews to uphold best practices and data quality standards.
- Represent MBZUAI at industry and research forums, showcasing technical capabilities in large-scale data processing and AI data infrastructure.
- Bachelor's degree in Computer Science, Data Science, Engineering, or a related technical field required
- Master’s degree or PhD degree or equivalent experience in Computer Science, Data Engineering, or related technical fields preferred.
- Extensive experience in data engineering, data processing, and automation using Python.
- Demonstrated proficiency in designing and deploying web crawling solutions, automated data extraction, and processing pipelines.
- Strong understanding of data structures, algorithms, databases, SQL, and performance optimization.
- Experience working with cloud infrastructure and distributed data processing frameworks (e.g., AWS, Spark, Kafka, Kubernetes).
- Excellent problem-solving abilities, attention to detail, and the capability to rapidly address technical challenges.
- Strong communication and collaboration skills with cross-functional teams.
- Proven track record of supporting NLP or AI research teams with rapid and reliable data delivery.
- Experience working with large language models, including evaluation, efficient inference, and prompt engineering.
- Experience with refining outputs from large-scale AI models, such as LLM-generated data.
- Contributions to open-source projects, coding competitions, or high visibility in coding communities (e.g., GitHub, Stack Overflow).
- Familiarity with the latest advancements in NLP data processing and large language model technologies.
About Ifm-us
Verified company details for this employer are not available yet.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Ifm-us
Distributed Machine Learning Engineer
Sunnyvale, USA
Discovered 1 weeks agoFull-time
Research Scientist - NLP
Sunnyvale, USA
Discovered 1 weeks agoFull-time
Machine Learning Engineer
Sunnyvale, USA
Discovered 1 weeks agoFull-time
Full Stack Software Engineer
Sunnyvale, USA
Discovered 1 weeks agoFull-time
Research Scientist - Data
Sunnyvale, USA
Discovered 1 weeks agoFull-time
AIE Internship
Sunnyvale, USA
Discovered 1 weeks agoInternship
Machine Learning Infrastructure Engineer
Sunnyvale, USA
Discovered 1 weeks agoFull-time
Senior Product Manager
Sunnyvale, USA
Discovered 1 weeks agoFull-time