Research Scientist, Web Data
Is this job for you?
Build my CV Build your CV and see how well you match this role — and every other one.
The role
Own and lead improvements to meaningful chunks of the web data pipeline, including scraping raw HTML into clean text data and integrating relevant image data.
Develop and apply measurements of weakness, via model evals or data pipeline statistics, to drive progress.
Define a medium-term agenda to improve data quality and pipeline efficiency, and build consensus with peers and stakeholders.
Collaborate with partner teams to leverage existing solutions and communicate necessary infrastructure improvements.
Execute day-to-day work through coding, running experiments, and reviewing contributions.
Qualifications include at least 3 years of self-directed work, building large-scale data pipelines (>=100M examples) in Python and/or C++, and evaluating pretrained LLMs.
Develop and apply measurements of weakness, via model evals or data pipeline statistics, to drive progress.
Define a medium-term agenda to improve data quality and pipeline efficiency, and build consensus with peers and stakeholders.
Collaborate with partner teams to leverage existing solutions and communicate necessary infrastructure improvements.
Execute day-to-day work through coding, running experiments, and reviewing contributions.
Qualifications include at least 3 years of self-directed work, building large-scale data pipelines (>=100M examples) in Python and/or C++, and evaluating pretrained LLMs.
See the full job post
Responsibilities, requirements, skills and benefits — create your free account.
Already have an account? Log in
Similar openings
Other roles that could suit you.
Remote workno
CityLondon, United Kingdom