
- מיקום
- כל הארץ
- היקף משרה
- משרה מלאה
- תפקיד
- דאטה סיינטיסט
תיאור המשרה
Key Responsibilities:
- Read research papers and source code critically. Reconstruct the mechanism, check its assumptions and complexity, and turn uncertain claims into testable hypotheses.
- Implement and modify model components, training loops and inference paths in Python and PyTorch. Build reference implementations and verify mathematical and numerical correctness before optimizing.
- Own experiments end to end: prepare data and environments, launch and monitor GPU runs, recover failed jobs, and preserve configurations, checkpoints and results so others can reproduce them.
- Design matched baselines and ablations. Test quality and failure modes across model sizes, datasets and context lengths, including retrieval, reasoning and generalization.
- Measure memory use, data movement, latency and throughput. Profile prefill and decoding separately, account for routing and kernel overhead, and check whether theoretical savings produce measurable system gains.
- Diagnose failures across mathematics, architecture, data and implementation. Work with researchers and compiler, runtime and hardware engineers to improve designs and integrate validated components.
Requirements:
- 3 to 5 years of hands-on AI/ML engineering or applied research experience, with evidence of
- implementing, running and stress-testing model work under real compute constraints.
- Strong foundations in linear algebra, probability, optimization and numerical methods. You can
- work through equations, follow tensor shapes and gradients, and connect a mathematical
- formulation to its implementation.
- Deep understanding of Transformer internals, attention and autoregressive inference, plus
- familiarity with recurrent, state-space or hybrid sequence models. You can reason about how a
- model stores, updates and retrieves information.
- Practical depth in at least two area of efficient AI, such as memory-efficient language models and
- long-context inference, including KV-cache compression, sparse or linear attention, model
- memory and compression, quantization, conditional computation, or inference optimization, and
- the ability to learn adjacent approaches deeply.
- Strong Python and PyTorch skills, including custom modules, training and inference code, testing
- and debugging. You can understand and repair code written by others or generated with AI
- assistance.
- Experience running GPU experiments in Linux environments, using version control, profiling
- workloads and managing reproducible results. You understand memory capacity, bandwidth,
- parallelism and numerical precision.
- Sound experimental judgment. You can distinguish a promising small-scale result from evidence
- that a method generalizes, identify confounded comparisons, and explain negative results and
- remaining uncertainty.
- The ability to work from an incomplete specification, choose the next useful experiment and
- communicate findings clearly to both researchers and engineers
Additional experience we value
• C++, CUDA or Triton; custom GPU kernels; distributed training; or work with compilers, inference runtimes and hardware-aware algorithm design.
• Model distillation, architecture conversion, fine-tuning or pretraining; scaling prototypes to larger models; or research implementations and open-source contributions that others have used.