Ogni settimana annunci remote-friendly, sia ibridi che full-remote, e consigli di carriera per muoverti meglio nel mercato tech e digital in Italia.
🏢 vCluster Labs
La posizione è stata chiusa oppure l'azienda non accetta più candidature.
Build and own the inference layer and production LLM serving pipeline at vCluster Labs. Deploy models to production on GPU infrastructure, operate serving frameworks (vLLM, SGLang, TensorRT-LLM), and optimize for scale using quantization, batching, caching, and routing. Implement real infrastructure in Python or Golang, partner with the CTO and Product to set the roadmap, and ensure low-latency, cost-effective serving. Required: production LLM serving experience, hands-on inference optimization, strong Python or Golang engineering skills, and clear communication. Bonus: Docker/Kubernetes, PyTorch/Transformers, CUDA/NCCL, DGX and NVIDIA Dynamo experience.
Livello di esperienza
Senior
Modalità di lavoro
Full-Remote
Retribuzione annuale
190.000€ - 230.000€