
Research Team Lead – Distributed AI Systems & Large-Scale Infrastructure 339
עלתה ל-JOBTIME לפני 14 שעות· בתוקף עד 24 בנובמבר 2026
- מיקום
- כל הארץ
- היקף משרה
- משרה מלאה
- תפקיד
- ראש צוות
תיאור המשרה
Requirements
- B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field
- 8+ years of experience in systems software, distributed computing, or AI infrastructure, with 3+ years in a leadership or team lead role
- Deep expertise in large-scale communication systems: collective communication, RDMA, network topology-aware routing, and bandwidth optimization
- Hands-on experience building software infrastructure for distributed training on custom accelerators or heterogeneous hardware %28GPU, NPU, TPU%29
- Strong knowledge of runtime systems: scheduling, execution graphs, kernel dispatch, synchronization primitives, and pipeline management
- Experience with memory management at scale: activation checkpointing, tensor offloading, rematerialization, KV cache management
- Proficiency in C/C++ and Python, with a focus on high-performance, production-quality code in Linux environments
- Proven ability to define technical vision, lead multi-person projects end-to-end, and deliver results under research and engineering timelines
- Excellent communication skills in English — confident presenting to international audiences, writing technical reports, and driving cross-team alignment
- Strong collaborative mindset and experience working in globally distributed, multicultural teams
Ways to Stand Out From the Crowd
M.Sc. or Ph.D. in a relevant field, with a strong publication record at systems or ML venues %28EuroSys, OSDI, SC, NeurIPS, MLSys, ISCA%29
- Hands-on experience with communication frameworks such as NCCL, MPI, HCCL, or UCX
- Experience with compiler and graph optimization for AI workloads %28XLA, TVM, Triton, or custom operator fusion%29
- Background in mixed-precision training, model parallelism %28Tensor Parallelism, Pipeline Parallelism, Expert Parallelism%29, and large model co-design
- Experience profiling and debugging performance bottlenecks on heterogeneous clusters using tools like Chrome tracing, nsight, or custom profilers