Research Engineer, Data Ingestion
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
Track your applications on mobile The free Whileresume app, on iPhone and Android.
The role
Support the design, implementation, and maintenance of a large-scale web crawler and data ingestion pipelines.
Design and run experiments to evaluate data quality, extraction methods, and crawling strategies, and analyze results to identify improvements.
Analyze crawled data to uncover patterns, gaps, and opportunities for data enhancement and source diversification.
Build pipelines for ingestion, analysis, and ongoing quality improvement, and develop specialized crawlers for high-value data sources.
Collaborate with Pretraining and Tokens teams to create feedback loops between crawled data and evaluation outputs; participate in code reviews and debugging.
Hybrid, office-based role in San Francisco with at least 25% in-person time; balance hands-on engineering with data research at scale.
Design and run experiments to evaluate data quality, extraction methods, and crawling strategies, and analyze results to identify improvements.
Analyze crawled data to uncover patterns, gaps, and opportunities for data enhancement and source diversification.
Build pipelines for ingestion, analysis, and ongoing quality improvement, and develop specialized crawlers for high-value data sources.
Collaborate with Pretraining and Tokens teams to create feedback loops between crawled data and evaluation outputs; participate in code reviews and debugging.
Hybrid, office-based role in San Francisco with at least 25% in-person time; balance hands-on engineering with data research at scale.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
or
Already have an account?
Log inYou might also like these jobs
No closely matching jobs yet — here are the most recent ones.