
Senior / Lead GPU Software Engineer %28AI Inference & High-Performance Compute%29
עלתה ל-JOBTIME לפני 6 שעות· בתוקף עד 24 בנובמבר 2026
- מיקום
- כל הארץ
- היקף משרה
- משרה מלאה
- תפקיד
- מהנדס/ת תוכנה
תיאור המשרה
Key Responsibilities
- High-Performance Kernel Development: Design and develop performance-critical software optimized for modern GPU architectures, writing custom GPU kernels and parallel computational workloads.
- Full-Stack Performance Optimization: Identify CPU/GPU bottlenecks to optimize end-to-end system performance, drastically improving hardware utilization, memory efficiency, throughput, and latency.
- Distributed Multi-GPU Scale: Design and scale architectures for GPU-intensive workloads, managing memory movement, synchronization, parallelism, and compute efficiency across multi-GPU and distributed computing environments.
- Profiling & Deep System Triage: Analyze complex system performance using advanced profiling and benchmarking tools to investigate difficult hardware-software interface problems and develop practical, highly optimized solutions.
- Technical Project Leadership: Lead technically complex projects, drive core architectural decisions, mentor engineers, and establish best practices for GPU and performance-oriented development.
- Ecosystem Evaluation: Continually evaluate new GPU architectures, open-source compilation libraries, deep learning frameworks, and acceleration technologies to maintain our competitive edge.
Key Requirements & Qualifications
Required Technical Skills
- Education & Experience: Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field, with 5+ years of software engineering experience and significant hands-on depth in GPU computing or high-performance computing %28HPC%29.
- Expert GPU Programming: Deep engineering mastery of CUDA and parallel programming methodologies.
- Systems-Level C++: Excellent programming skills in modern C++ with a proven track record of developing performance-critical or low-level software.
- Architectural Mastery: Firm structural understanding of GPU architecture, parallel programming patterns, memory hierarchies, multithreading, and synchronization primitives.
- Data Path Competence: Strong grasp of CPU/GPU interaction, PCIe data movement, and asynchronous compute streams.
- Production-Grade Delivery: Demonstrated ability to independently investigate complex system anomalies, isolate root causes, and take intricate software blocks from design to production.
Preferred Qualifications & Domain Expertise
- NVIDIA Tooling Mastery: Expert familiarity with NVIDIA GPU architectures and profiling tooling %28e.g., Nsight Systems, Nsight Compute, or similar deep-dive profilers%29.
- AI/ML Infrastructure Integration: Direct experience with AI/ML inference or training infrastructure, optimizing latency and throughput for real-world models.
- Advanced Frameworks & Libraries: Practical familiarity with modern deep learning ecosystems and communication primitives, including:
- Frameworks & Compilers: PyTorch, Triton, and MLIR.
- Acceleration Libraries: NCCL, cuBLAS, and cuDNN.
- Data Center Networking & Scale: Understanding of networking technologies relevant to distributed GPU clustering, including RDMA, InfiniBand, or RoCE.
- Linux Systems Internals: Strong experience with Linux systems and low-level performance analysis tools.
- Leadership Background: Previous technical leadership or team-lead experience in an agile, high-scale engineering environment.