
- מיקום
- כל הארץ
- היקף משרה
- משרה מלאה
- תפקיד
- מהנדס/ת תוכנה
תיאור המשרה
Description:
seeking an exceptional, hands-on Senior GPU & VLIW Compiler Engineer to design, optimize, and build the compiler graph transformations and lower-level code generation paths for our AI inference acceleration solutions. In this role, you will be a core technical driver bridging high-level machine learning frameworks down to peak machine-utilization machine code for both modern parallel GPU pipelines and custom VLIW %28Very Long Instruction Word%29 execution cores.
Your mandate will focus on building high-performance compilation layers, custom target backends, and optimization passes that map dense tensor pipelines and large language models %28LLMs%29 to hardware efficiently. Working in a tight-knit startup environment, you will use a blend of systems-level C++ and fast-prototyping Python environments to collaborate closely with kernel engineers, deep learning software architects, and silicon designers to squeeze out maximum execution throughput, ultra-low latency, and memory locality.
Key Responsibilities:
- Compiler Frontend & Middle-End Optimization: Develop and implement high-performance graph lowering, schedule optimizations, and target-independent transformations within compilation frameworks to support state-of-the-art AI topologies.
- VLIW & GPU Backend Code Generation: Write target-specific backends and optimization passes focusing on memory placement, register allocation, auto-vectorization, and machine-level resource binding.
- Instruction Scheduling for VLIW: Architect complex compiler passes specifically optimized for VLIW processing blocks, mastering structural challenges like software pipelining, explicit loop tiling, instruction bundling, hazard detection, and static dependency scheduling.
- AI Ecosystem & Compiler Co-Design: Build infrastructure integrating modern machine learning compiler architectures, ensuring flawless translation loops through frameworks like Triton, MLIR, OpenXLA, TVM, and PyTorch.
- Python Infrastructure & Tooling: Drive the architectural integration of Python frontend tooling, utilizing Python to design compiler optimization layers, build automated code generation frameworks, implement graph-level heuristics, and automate deep verification pipelines.
- Multi-Chip & Hardware-Software Co-Design: Collaborate intensely with silicon architects and hardware engineering teams to evaluate workload bottlenecks. Use your compiler pipeline to model data movement across high-bandwidth interfaces %28PCIe, LPDDR6x, HBM, and UCIe die-to-die fabrics%29 to maximize compute cluster efficiency.
- Performance Analysis & Assembly Triage: Triage execution traces, analyze machine assembly output, and isolate compiler-induced performance drops on hardware simulators and prototypes.
Requirements:
- Education & Experience: Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field, with 5+ years of software engineering experience and a deep focus on compiler design, code generation, or high-performance computing %28HPC%29.
- Expert Programming Skills: Mastery of modern systems-level C++ for core compiler architecture construction, combined with strong proficiency in Python for compiler graph manipulation, scripting, and development tooling.
- Production Compiler Infrastructure: Practical engineering experience developing inside production compiler architectures %28such as LLVM/Clang backend pipelines, MLIR dialects, Triton, or OpenXLA%29.
- VLIW & GPU Architecture Domain Knowledge: Firm structural understanding of modern GPU microarchitectures %28SIMD/SIMT parallel execution models%29 and/or static parallel execution principles inherent to VLIW processors.
- Advanced Code Optimization Mastery: Solid grasp of assembly-level tuning patterns, instruction-level parallelism %28ILP%29, memory alignment rules, and explicit compiler scheduling algorithms.
- Startup Agility: Exceptional capability to thrive in a fast-paced, high-ambiguity startup environment. A pragmatic engineer who knows how to balance multi-pass compiler purity with rapid time-to-market performance wins.
Preferred Qualifications
- Deep Learning Compilation Exposure: Direct experience optimizing graph pipelines for modern generative AI and LLM inference deployment structures %28e.g., quantization mapping, flash-attention compilation%29.
- Hardware Bring-up & Emulation: Familiarity working with pre-silicon emulation environments or early prototype hardware drivers to validate early compiler codegen outputs.