Search by job, company or skills

Nvidia

Senior Manager, Cloud Services Platform

Save
  • Posted 22 days ago
  • Be among the first 10 applicants
Early Applicant

Job Description

For over 25 years, NVIDIA has been revolutionizing computer graphics, PC gaming, and accelerated computing. It's a unique legacy of innovation that's fueled by great technology-and amazing people.

At NVIDIA, we are seeking a highly skilled Senior Engineer Operations Manager to join our world-class NGC Cloud team. In this role, you will help drive the efficiency, reliability, and scalability of the systems that power our global business operations. This is an exceptional opportunity to shape how we automate, streamline, and support critical operational workflows across the organization. You will define how we implement innovative automation and support solutions, enabling teams to operate seamlessly and deliver impact at global scale-all within an encouraging and inclusive environment.

What you'll be doing:

  • Lead strategy, execution and operations for cloud services that provide container, artifact, and ML model registry capabilities for NVIDIA engineering teams.

  • Architect, Design, plan, implement and Operate complex PaaS for the GPU cloud services.

  • Partner with AI infrastructure, security, product, and engineering teams to define standards for model packaging, image build pipelines, artifact management, and deployment workflows.

  • Define and track KPIs and SLAs for registry and related services, including availability, latency, storage efficiency, reliability, and developer experience.

  • Drive operational excellence for secure, scalable cloud services

  • Mentor and coach engineering managers and senior individual contributors build a strong leadership bench and a healthy, inclusive engineering culture.

  • Influence architecture and technical direction while empowering teams to own detailed design and implementation decisions.

  • Communicate clearly with senior leadership on strategy, risks, execution progress, and outcomes represent the platform in multi-functional planning discussions.

What we need to see:

  • 10+ overall years of software engineering experience with significant ownership of cloud platforms, distributed systems, developer platforms, artifact registry systems, storage systems, or infrastructure services.

  • 5+ years of engineering leadership experience, including experience managing managers or leading multiple senior technical workstreams through other leaders.

  • Bachelors degree or equivalent experience.

  • Proven success operating customer-facing or company-critical services with demanding availability, latency, throughput, data integrity, and security expectations.

  • Strong technical judgment in registry, artifact management, or developer platform systems.

  • Deep cloud-native systems background across Kubernetes, object storage, relational or NoSQL databases, event streaming, caching, API design, service-to-service authentication, and observability.

  • Experience with containerization and registries such as Docker, Kubernetes, Docker Hub, Harbor, ECR, GCR, GAR, or similar technologies at enterprise scale.

  • Excellent people leadership skills, including hiring, performance management, career development, and building diverse, high-performing teams.

  • Strong communication skills, with experience presenting to senior executives and influencing cross-organization priorities.

Ways to stand out from the crowd:

  • Directly led teams building and operating distributed cloud platforms at enterprise scale.

  • You have led reliability transformations for large services, including SLO adoption, forecasting resource needs, load testing, and stress testing.

  • Worked in AI infrastructure, accelerated computing, model distribution, or regulated enterprise software delivery.

  • Track record of growing managers and senior technical leaders who can own ambiguous, high-impact platform areas without constant blocking issue.

Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family

More Info

About Company

NVIDIA&#8217&#x3B;s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI &#8212&#x3B; the next era of computing &#8212&#x3B; with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as "the AI computing company.

Job ID: 150474821

Similar Jobs

Bengaluru, India

Skills:

object storage Storage SystemsAPI designDockerDistributed SystemsInfrastructure ServicesKubernetesdeveloper platformsHarborNoSQL databasescloud platformsevent streamingservice-to-service authenticationartifact registry systemsCachingECRGCRGARDocker Hubobservability

Bengaluru, India

Skills:

Cloud TechnologiesAzureGoogle CloudAWSData AIAgile Delivery ModelsSaaS PlatformsMainframe ModernizationAPI-led IntegrationAiMicroservices Architectures

Bengaluru, India

Skills:

Security ComplianceVpcStorageNosqlTerraformDockerRESTful API developmentPythonAWSBashSqlSdlcJenkinsGcpIamPuppetHelmAzureKubernetesVNetagentic AIChefAI productivity toolsAI ML ecosystemsGoGitHub Actionslarge language modelsobservabilityComputeIstioAPI gatewaysArgoCD

Bengaluru, India

Skills:

bedrock ReactDockerAzure Data LakeFlaskAzure cloud servicesPythonAWSPower BiDatabricksFastAPIKubernetesAws S3knowledge graphsPineconeClaudeAzure OpenAIGPTLangGraphAWS SagemakerNext.jsMistralLLM frameworksChromaDBLangChainAnthropicSnowflake CortexGen AIFAISSRAGWeaviate

Bengaluru, India

Skills:

GolangReactTypescripttest automationKubernetesAWSdistributed systems testingCI CD