Jobherder
  • How it works
  • Pricing
  • Sample output
  • Jobs
  • Blog
  • Help
Log inTry a free sample

← All jobs

EU remote

System Engineer (Token Factory)

NebiusRemote; Remote - EuropePosted 2 Aug 2026

Start a search — €9.99Apply on employer site

About this role

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. About the role: Token Factory is a part of Nebius Cloud, one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale.

Responsibilities: Develop and optimize low-level kernels and runtime components for AI inference Improve performance of inference engines GPU platforms Profile and debug system-level and hardware-level performance issues Integrate support for new hardware architectures (Hopper, Blackwell , Rubin ) Collaborate with ML and backend teams to optimize end-to-end execution Required Qualifications: Strong proficiency in C++ , OR expertise in GPU programming with a focus on low-level high-performance coding and memory management Experience in GPU programming or systems-level software development , e.g.

operating system internals, kernel modules, or device drivers Hands-on experience with profiling and debugging tools to identify performance issues on both CPUs and GPUs, and the ability to optimize code based on those findings. Solid understanding of CPU/GPU architecture and memory hierarchy Preferred Qualifications: Experience with GPU computing programming : CUDA, ROCm , CUTLASS, Cute, ThunderKittens , Triton, Pallas, Mosaic GPU Familiarity with ML inference runtimes (e.g.

Jobherder

Stop searching. Start applying.

Product

How a search worksPlans and pricingExample deliveryJob board

Resources

ArticlesHelp centreFor recruitersAgentsAffiliate programme

Features

CV tailoringRemote job searchCareer change

TensorRT , TVM) Knowledge of Linux internals, drivers, or compiler toolchains <span data-

Source listing: greenhouse_nebius

Prefer jobs chosen for you?

Upload your CV and Jobherder returns handpicked roles with a tailored CV and cover letter for each — built from your real experience.

Start a search — €9.99

© 2026 Jobherder

PrivacyTermsSupport