
- מיקום
- כל הארץ
- היקף משרה
- משרה מלאה
- תפקיד
- דאטה סיינטיסט
תיאור המשרה
Role Overview
We are looking for a Senior/Staff AI Algorithms Engineer to join our research and engineering team, focusing on the development, training, and optimization of large-scale language models %28LLMs%29 on distributed networks. The ideal candidate combines deep theoretical understanding with hands-on engineering excellence.
Must-Have Requirements
Education & Experience
- M.Sc. or Ph.D. in Computer Science, Electrical Engineering, Applied Mathematics, or a related field
- 5+ years of industry or research experience in machine learning / deep learning
- Demonstrated track record of delivering production-quality ML systems or publishing in top-tier venues %28NeurIPS, ICML, ICLR, TMLR, etc.%29
Core AI/ML Expertise
- Deep understanding of transformer architectures %28encoder-only, decoder-only, encoder-decoder%29, attention mechanisms, positional encodings %28RoPE, ALiBi, etc.%29, and normalization strategies
- Hands-on experience training large-scale models %281B–70B+ parameters%29 from scratch
- Familiarity with pre-training, instruction tuning, RLHF, DPO, and related alignment techniques
- Knowledge of model evaluation: perplexity, downstream benchmarks %28MMLU, HellaSwag, etc.%29, and ablation methodology
Distributed Training & Systems
- Strong practical experience with distributed training paradigms:
- Data Parallelism %28DDP, FSDP%29
- Tensor Parallelism %28Megatron-style%29
- Pipeline Parallelism
- ZeRO %28Stage 1/2/3 with optimizer/gradient/parameter sharding%29
- Proficiency with modern training frameworks: PyTorch, DeepSpeed, Megatron-LM, Hugging Face Accelerate / Transformers
- Experience managing large-scale GPU clusters %28A100/H100/B200 or equivalent%29, including job scheduling, multi-node communication %28NCCL%29, and GPU utilization monitoring
Engineering Skills
- Expert-level Python programming; clean, testable, modular code
- Proficiency with data pipelines for LLM pre-training
- Solid understanding of profiling and debugging training runs: loss spikes, gradient norms, throughput bottlenecks %28MFU%29, dead nodes
Strong Advantage %28Nice-to-Have%29
Research & Innovation
- First/co-author publications in LLM training, efficient transformers, or distributed ML at top venues
- Experience with novel architecture exploration: SSMs %28Mamba%29, MoE, hybrid architectures
- Familiarity with continual learning or domain adaptation
- Experience with federated learning or layer-wise / alternative training strategies