
AI Engineer
Nvidia
About this job
NVIDIA, a world leader in accelerated computing and AI, is actively seeking highly skilled and passionate AI Systems Engineers to join our innovative team in Pune, India. This pivotal role offers a unique opportunity to contribute to the future of artificial intelligence by developing and optimizing high-performance AI inference runtimes for NVIDIA’s cutting-edge GPU platforms. You will be instrumental in driving advancements across Generative AI technologies, impacting a wide range of applications on RTX, RTX Pro, and DGX GPUs.
As an AI Systems Engineer, you will be at the heart of developing solutions that enable faster, more efficient, and scalable AI models, working with some of the most advanced hardware and software in the industry. Your expertise will directly influence the performance and capabilities of AI systems used globally.
Key Responsibilities
- Develop, optimize, and enhance AI inference runtimes specifically tailored for NVIDIA's state-of-the-art RTX, RTX Pro, and DGX GPUs, ensuring maximum performance, efficiency, and robustness.
- Contribute significantly to the full development lifecycle across various critical AI frameworks and platforms, including Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX.
- Focus on meticulous optimization of performance, memory utilization, scalability, and accuracy for a diverse range of Generative AI (GenAI) workloads, such as Large Language Models (LLMs), Vision-Language Models (VLMs), Text-to-Speech (TTS), Automatic Speech Recognition (ASR), and Diffusion models.
- Actively collaborate with cross-functional teams, including software development, research, architecture, and product management, to strategically shape the future and widespread adoption of local AI technologies.
- Participate in code reviews, contribute to architectural discussions, and mentor junior engineers, fostering a culture of technical excellence and continuous improvement.
Requirements
- A minimum of 7 years of hands-on industry experience in AI/ML software development, with a strong preference for candidates in senior engineering roles demonstrating leadership potential.
- Exceptional proficiency in C++, coupled with a robust understanding of data structures, algorithms, and advanced debugging techniques for complex systems.
- Demonstrated experience in developing and optimizing AI inference pipelines and working effectively with various Machine Learning (ML) and Deep Learning (DL) frameworks.
- Deep technical understanding of runtime internals, including KV-cache mechanisms, efficient scheduling, sophisticated memory management strategies, and quantization techniques for AI models.
- Practical, hands-on experience with CUDA programming, GPU architecture, performance optimization methodologies, and contributions to open-source AI frameworks will be considered a significant advantage.
- A strong academic background from top-tier institutions is highly preferred, demonstrating a solid foundation in computer science, electrical engineering, or a related quantitative field.
What We Offer
- The unparalleled opportunity to work at the cutting edge of AI and GPU technology with a global leader, contributing to products and research that define the future of computing.
- A collaborative, innovative, and fast-paced work environment where your ideas are valued, and you can make a tangible impact on groundbreaking AI systems.
- Competitive compensation packages, comprehensive health and wellness benefits, and continuous opportunities for professional growth and skill development.
- Exposure to diverse and challenging projects, working alongside some of the brightest minds in the industry on complex global problems.
- A vibrant company culture that encourages technical excellence, creativity, and a healthy work-life balance, supporting your overall well-being.
Eligibility
Professionals • 7-7 years of experience