Remote
Web Scraping Specialist
About this role
Who We Are: We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models. We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.
We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI. The Role. We are seeking a Web Scraping Specialist who is proficient and brings significant experience in data extraction and web scraping techniques. You will join a small, specialized team and lead efforts to gather and analyze data, optimize scraping processes, and support our vision for a future where Grass plays a crucial role in transforming internet data accessibility.
Please note: This role requires a work schedule that sufficiently overlaps with EST business hours to collaborate effectively with the team. Who You Are. - Demonstrated ability to extract data from complex websites with minimal supervision, with a portfolio or examples of past projects. - Proficiency in languages such as Python or JavaScript, with strong skills in libraries and frameworks like BeautifulSoup, Scrapy, or Selenium.
- Knowledge of asynchronous programming, multithreading, and distributed scraping. - In-depth knowledge of HTML, CSS, JavaScript, and the Document Object Model (DOM). - Experience with NoSQL databases (MongoDB, Cassandra), capable of designing efficient storage solutions and managing data integrity. - Ability to apply machine learning algorithms for data cleaning, categorization, or predictive analysis adds significant value.
- Experience with cloud services (AWS, Google Cloud, Azure) for deploying and managing scraping jobs at scale. - Active participation in open-source projects related to web scraping, data processing, or similar fields. What You'll Be Doing. - Write, test, and refine code that extracts data from various online sources, ensuring reliability and efficiency. - Perform data retrieval tasks, handling complexities such as pagination and dynamic content loaded with AJAX.