Disponibile su Google Play Entra nel Talent Radar
Newsletter settimanale

TechCompenso per Te

Ogni settimana annunci remote-friendly, sia ibridi che full-remote, e consigli di carriera per muoverti meglio nel mercato tech e digital in Italia.

V

Senior Inference Engineer

🏢 vCluster Labs

Full-Remote
⚠️

La posizione è stata chiusa oppure l'azienda non accetta più candidature.

📝 Descrizione

Build and own the inference layer and production LLM serving pipeline at vCluster Labs. Deploy models to production on GPU infrastructure, operate serving frameworks (vLLM, SGLang, TensorRT-LLM), and optimize for scale using quantization, batching, caching, and routing. Implement real infrastructure in Python or Golang, partner with the CTO and Product to set the roadmap, and ensure low-latency, cost-effective serving. Required: production LLM serving experience, hands-on inference optimization, strong Python or Golang engineering skills, and clear communication. Bonus: Docker/Kubernetes, PyTorch/Transformers, CUDA/NCCL, DGX and NVIDIA Dynamo experience.

🔎 Informazioni

💼

Livello di esperienza

Senior

🖥️

Modalità di lavoro

Full-Remote

💰

Retribuzione annuale

190.000€ - 230.000€

🔹 LLM 🔹 vLLM 🔹 SGLang 🔹 TensorRT-LLM 🔹 quantization 🔹 batching 🔹 caching 🔹 routing 🔹 Python 🔹 Golang 🔹 Docker 🔹 Kubernetes 🔹 PyTorch 🔹 Transformers 🔹 CUDA 🔹 NCCL 🔹 NVIDIA Dynamo 🔹 GPU 🔹 DGX