NVIDIA logo

Senior System Software Engineer, Speech AI

NVIDIA · Posted Sep 30

Graphics processing units, artificial intelligence platforms, accelerated computing, and data center solutions

India, PuneFull-timeOnsiteSenior Level6+ years₹40.0L–₹70.0L yearly100+ applicants
SemiconductorsArtificial IntelligenceComputer HardwareCloud ComputingPublic Company
Full time

Get a personal compatibility score

Add a resume for personal matches

About the role

NVIDIA's Speech AI Engineering team builds GPU-accelerated AI systems and supports enterprise and developer customers deploying conversational AI. The team focuses on Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Audio Language Models (ALM), and Speech-to-Speech (S2S) pipelines running at scale. This senior role involves hands-on customer engagement to implement, troubleshoot, and optimize speech systems in production, along with building and optimizing C++ and Python backend implementations and client SDKs.

What you will do

  • Work on cutting-edge GPU-accelerated AI systems deployed at scale
  • Tackle challenging problems in real-time streaming audio processing and low-latency inference
  • Troubleshoot and resolve complex issues across ASR, TTS, ALM, and S2S pipelines
  • Model Integration: Work alongside Model researchers to transition ASR, TTS and S2S models from research to production readiness.
  • Develop Core Speech Services: Build and enhance C++ & python backend implementations for ASR, TTS, and S2S pipelines, leveraging CUDA for GPU acceleration
  • Optimize Inference Performance: Improve streaming latency and throughput through advanced batching strategies, encoder caching, and multi-threaded pipeline optimizations
  • Feature Development: Add new capabilities such as advanced voice activity detection, speaker diarization, decoder implementations (CTC, WFST, Flashlight), and text post-processing
  • Client Libraries: Contribute to Python and C++ client SDKs and CLI tools for easy service integration
  • Assist with API integration, SDK usage, model deployment, and performance optimization
  • Provide advanced technical guidance to customers implementing speech technology solutions

Skills used in this role

C++PythonCUDAGPUHTTP RESTgRPCWebSocketsTransformersCNNsRNNsLLMCTCWFSTFlashlightAutomatic Speech Recognition (ASR)Text-to-Speech (TTS)Audio Language Models (ALM)Speech-to-Speech (S2S)Speaker DiarizationVoice Activity DetectionDebuggingCollaboration

What the employer is looking for

  • Masters or BE/BTech in Computer Science, computer architecture, or related field
  • 6+ years of experience
  • Excellent C++ & Python programming and software design skills, including debugging, performance analysis, and test design.
  • Experience with inference pipelines for LLM, Speech Recognition & Speech Synthesis
  • Solid understanding of modern model architectures (Transformers, CNNs, RNNs)
  • Excellent debugging abilities spanning multiple software (storage systems, kernels and containers)
  • Experience building and deploying cloud services using HTTP REST, gRPC, Websockets and related technologies
  • Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment
  • Ability to work independently, define project goals and scope and manage your own development effort.
  • Knowledge of real-time streaming audio systems and low-latency architectures
  • Experience with speech model fine-tuning or customization

Preferred qualifications

  • Publications or contributions to ML optimization open-source projects.
  • Experience with embedded systems or edge deployment

Benefits and support

  • Compensation and benefits are detailed in the job posting

About NVIDIA

NVIDIA Corporation is a multinational technology company that designs graphics processing units (GPUs) for the gaming and professional markets, as well as system on a chip units (SoCs) for the mobile computing and automotive market. Pioneering accelerated computing, the company has become a driving engine of modern artificial intelligence, deep learning, and data center infrastructure.

Industry
Semiconductors
Company size
42000+ employees
Founded
April 5, 1993
Location
Santa Clara, California, USA
Funding stage
Public Company

Funding

Public Company · $5M raised

Worldview Technology Partners

Leadership

JH
Jensen Huang

Founder, President and Chief Executive Officer

CK
Colette Kress

Executive Vice President and Chief Financial Officer

CA
Chris A. Malachowsky

Founder and NVIDIA Fellow

JP
Jay Puri

Executive Vice President, Worldwide Field Operations