Remote
Machine Learning Engineer
About this role
Who We Are: We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models. We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.
We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI. The Role: We are looking for a Machine Learning Engineer with strong skills and significant experience developing machine learning models. You will join a small, innovative team and lead efforts to advance our capabilities, drive model development, and support our vision for a future where Grass is transformative in the internet's evolution.
Please note: This role requires a work schedule that sufficiently overlaps with EST business hours to collaborate effectively with the team. Who You Are: - Bachelor’s, Master’s, or Doctoral degree in Data Science, Computer Science, Statistics, or a related field. - A minimum of 3 years of work or research experience dealing with large datasets. - Experience working with large-scale text datasets, NLP pipelines, or data preparation for LLM training is highly preferred.
- Strong coding skills in Python or other object-oriented programming languages. - Experience with text deduplication, dataset filtering, corpus curation, or data distillation is a strong plus. - Graduate-level knowledge of statistics, including but not limited to hypothesis testing, regression analysis, and probability. - Excellent work ethic and the ability to thrive in a fast-paced startup environment. - Strong problem-solving skills and attention to detail.
- Good communication skills, with the ability to articulate complex data concepts to non-technical stakeholders. - Experience working in a high-output team. What You'll Be Doing: - Developing data processing pipelines and machine learning solutions for large-scale NLP and LLM applications, including improving the quality, filtering, and preparation of training datasets. - Designing and implementing pipelines for processing and analyzing large datasets.