EU remote
Senior AI Engineer, Unstructured AI
About this role
Joining Collibra’s Unstructured AI Team Work at the forefront of context engineering - shaping how AI systems retrieve, structure, and leverage context to deliver accurate, high-quality results at scale. Own end-to-end technical delivery of Unstructured AI systems - from feature prototype to stable production across enterprise environments. Build and scale full-stack systems that ingest, process, and enrich large volumes of unstructured content from distributed enterprise silos (PDFs, contracts, reports, and other document types).
Collaborate with the Best: Work closely with xYC Founders to understand complex business challenges and deliver Deasy to solve them. Be part of a dynamic team where ideas flow freely and creativity thrives. Learn and Lead: Stay ahead of the curve by engaging with the latest developments in machine learning and AI. Share knowledge and lead by example to maintain high building standards. This is a hybrid role based in our Brussels office.
Our hybrid model means you’ll work from the office at least two days each week. This setup helps us stay connected, work more closely together, and keep making progress as a team. Senior AI Engineer at Collibra are responsible for Shipping complex systems under real ambiguity — often defining the scope and acceptance criteria yourself, not receiving them. Writing and reviewing production-grade backend code (Python, FastAPI) — you write the services you design, rather than handing implementation to someone else.
Building/deploying document-processing systems that handle large-scale, unstructured data environments. Integrating data from diverse enterprise data sources (e.g., SharePoint, Salesforce, or internal APIs) to provide context for AI features. Partnering across engineering, product, and sales teams, ensuring alignment from prototype to rollout. Occasionally working with modern frontend development. You have Strong proficiency in Python (data processing, API development, and integrations).
Hands-on work with LLM-based and AI-driven enrichment models (e.g., classification, entity extraction, deduplication, PII detection). Production experience with Spark or comparable big data frameworks — you've tuned and debugged jobs at real scale, not just written ones that worked on sample data. Experience shipping tested, reviewed production services rather than notebooks — and the discipline to hold that line when a coding agent writes the first draft.