
- מיקום
- כל הארץ
- היקף משרה
- משרה מלאה
- תפקיד
- דאטה סיינטיסט
תיאור המשרה
Responsibilities:
- Build and maintain end-to-end integration flows across the AI inference pipeline %28serving, orchestration, APIs, and infrastructure%29
- Design, implement, and optimize LLM inference workflows, including prefill and decode stages
- Improve system performance with focus on throughput, latency, and interactivity
- Write production-grade components in Python and integrate them into the broader system
- Contribute to system-level logic such as smart hardware selection and execution strategies
- Integrate models %28open source and custom%29, services, and APIs into cohesive, reliable end-to-end application pipelines
Requirements
- 4+ years of experience in software engineering or machine learning engineering
- Strong proficiency in Python
- Strong experience with LLM inference systems and performance optimization
- Hands-on experience with system integration and end-to-end workflows
- Experience with inference frameworks such as vLLM, TensorRT, SGLang etc
- Experience working with GPU/accelerator-based systems
Preferred Qualifications
- Hands-on experience with Dynamo and LLM-D for LLM inference and serving
- Familiarity with Kubernetes and cloud environments