TixelJobs
N
Nebiusvia Greenhouse

Senior Applied Scientist, Efficient LLM Inference & Model Optimization

Palo Alto, California, United StatesPosted 5d ago
NLP / LLMSeniorFull-time

Not sure if you're a good fit?

Upload your resume and TixelJobs AI will compare it against Senior Applied Scientist, Efficient LLM Inference & Model Optimization at Nebius. Get a match score, missing keywords, and improvement tips before you apply.

Free preview · Your resume stays private

About the Role

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role 

Nebius Token Factory needs scientists who can turn frontier inference bottlenecks into research problems, publish credible work, and then help ship the results into production. This is not a papers-only research role. The Applied Scientist is expected to design rigorous experiments, write strong code, collaborate with engineers, and convert research into deployed inference capabilities.

A Senior Applied Scientist owns well-scoped research and production optimization projects. They can publish or prepare high-quality technical work while also producing code, experiments, and prototypes that engineers can use.

Your responsibilities: 

  • Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff.

  • Prepare internal reports, technical blogs, or papers when the work is externally credible.

  • Partner directly with MLEs to ensure research prototypes become usable production components.

  • Define and execute research programs in efficient LLM and VLM inference with measurable production impact.

  • Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse, KV-cache compression, long-context inference, MoE routing, and model/runtime co-optimization.

  • Build high-quality prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks, then work with MLEs and platform engineers to productionize them.

  • Design rigorous evaluation methodology covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.

  • Publish papers, technical reports, blog posts, and open-source artifacts that build external credibility for Nebius Token Factory.

  • Collaborate with MLE, GPU kernel, backend infrastructure, product, and customer teams to choose high-leverage research bets.

  • Mentor engineers and scientists on experimental design, scientific rigor, and model/system tradeoffs.

Must-haves: 

  • PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field.

  • Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.

  • Strong hands-on coding ability in Python and PyTorch; ability to move from idea to experiment to prototype quickly.

  • Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs.

  • Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis.

  • Excellent written and verbal communication.

Nice-to-haves: 

  • First-author publications in NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, HPCA, or comparable venues.

  • Experience deploying ML models or inference optimizations in production.

  • Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals.

  • Experience with post-training,

    Share