
AI Systems Engineer
Nvidia
About this job
NVIDIA, the pioneer of GPU-accelerated computing and a leader in artificial intelligence, is actively seeking highly skilled and passionate AI Systems Engineers to join our dynamic team in Pune, India. This is an unparalleled opportunity to contribute to the cutting-edge of AI inference technology, working on solutions that power the future of local AI on NVIDIA's renowned RTX, RTX Pro, and DGX GPUs. If you thrive on solving complex performance challenges and have a deep understanding of AI runtime optimization, we invite you to help shape the next generation of intelligent systems.
As an AI Systems Engineer at NVIDIA, you will be instrumental in developing and optimizing critical AI inference runtimes. You will engage with advanced deep learning frameworks and low-level GPU programming to ensure our hardware delivers unparalleled performance and efficiency for a diverse range of Generative AI workloads. Your contributions will directly impact millions of users and developers, pushing the boundaries of what's possible with AI.
Key Responsibilities
- Architect, build, and meticulously optimize AI inference runtimes specifically tailored for NVIDIA's RTX, RTX Pro, and DGX GPUs.
- Develop and integrate across a spectrum of advanced AI frameworks and technologies, including Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX.
- Focus intensely on optimizing performance, memory utilization, scalability, and accuracy for various Generative AI workloads, encompassing Large Language Models (LLMs), Vision-Language Models (VLMs), Text-to-Speech (TTS), Automatic Speech Recognition (ASR), and Diffusion models.
- Engage in close collaboration with cross-functional software, research, architecture, and product teams to strategically influence and define the future direction of local AI solutions.
- Troubleshoot and debug complex issues within AI inference pipelines, ensuring robust and reliable system operation.
Requirements
- A minimum of 7 years of relevant industry experience in software development, with a strong emphasis on high-performance computing and AI systems. Senior professionals with extensive experience are particularly encouraged to apply.
- Exceptional proficiency in C++, complemented by robust skills in data structures, algorithms, and advanced debugging techniques.
- Demonstrated experience with AI inference pipelines and a solid understanding of various Machine Learning (ML) and Deep Learning (DL) frameworks.
- A deep and comprehensive understanding of runtime internals, including KV-cache mechanisms, efficient scheduling, sophisticated memory management strategies, and quantization techniques.
- Hands-on experience with CUDA and GPU programming, coupled with a proven track record in performance optimization.
- Familiarity and practical experience with open-source AI frameworks are considered a significant advantage.
- Strong academic background, preferably from a top-tier institution, demonstrating a solid foundation in computer science or a related field.
What We Offer
- The opportunity to work at the forefront of AI innovation with a global leader in GPU technology.
- A collaborative and intellectually stimulating environment where your contributions have a tangible impact on cutting-edge products.
- Exposure to a wide array of advanced AI models and frameworks, fostering continuous learning and skill development.
- Competitive compensation and the chance to grow your career within a company renowned for its technological breakthroughs.
- An inclusive culture that values diverse perspectives and encourages bold problem-solving.
Eligibility
Professionals • 7-7 years of experience