Jobs
>
Santa Clara

    Senior Production SRE Engineer - Santa Clara, United States - NVIDIA

    Default job background
    Description
    Senior Production SRE Engineer - Storage page is loaded

    Senior Production SRE Engineer - Storage

    Apply

    locations

    US, CA, Santa Clara

    US, Remote

    time type

    Full time

    posted on

    Posted 4 Days Ago

    job requisition id


    JR


    Site Reliability Engineering (SRE) is an engineering discipline that involves designing, building, and maintaining large-scale production systems with high efficiency and availability.

    It encompasses various areas, including software and systems engineering practices, storage, data management, and services.

    SRE professionals are highly specialized and possess expertise in different domains such as systems, networking, storage, coding, database management, capacity management, continuous delivery, and deployment, as well as open-source cloud-enabling technologies like Kubernetes, containers, and virtualization.

    Their responsibilities encompass ensuring reliable storage solutions, managing data efficiently, and providing related services to support the overall stability and performance of the production systems.

    SRE at NVIDIA ensures that our internal and external facing GPU cloud services have reliability and uptime as promised to the users and at the same time enables developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency, and performance.

    SRE is also a mindset and a set of engineering approaches to running better production systems and optimizations.

    Much of our software development focuses on eliminating manual work through automation, performance tuning, and growing the efficiency of production systems.

    As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems.

    Practices such as limiting time spent on reactive operational work, blameless postmortems, and proactive identification of potential outages factor into iterative improvement that is key to product quality and interesting and dynamic day-to-day work.

    SRE's culture of diversity, intellectual curiosity, problem-solving, and openness is important to its success. Our organization brings together people with a wide variety of backgrounds, experiences, and perspectives. We encourage them to collaborate, think big, and take risks in a blame-free environment.

    We promote self-direction to work on meaningful projects while striving to build an environment that provides the support and mentorship needed to learn and grow.


    What You Will Be Doing:
    Assist in the design, implementation, and support of large-scale storage clusters, including monitoring, logging, and alerting.


    Work with AI/ML workloads to capture and correlate behavior in large clusters and workflows, which are otherwise hard to understand.


    Work closely with peers on the team to improve the lifecycle of services – from inception and design, through deployment, operation, and refinement.


    Support services before they go live through activities such as system design consulting, developing software and frameworks, capacity management, and launch reviews.


    Maintain services once they are live by measuring and monitoring availability, latency, and overall system health, including leveraging machine learning models.


    Scale systems sustainably through mechanisms like AI/ML and automation, and evolve systems by pushing for changes that improve reliability and velocity.

    Practice sustainable incident response and blameless postmortems.

    Be part of an on-call rotation to support production systems.


    What We Need To See:
    BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics) or equivalent experience.

    At least 5+ years practical experience.

    Experience with algorithms, data structures, complexity analysis, software design, and maintaining large-scale Linux-based systems.

    Experience in one or more of the following: C/C++, Java, Python, Go, Perl or Ruby, AI/ML frameworks and methodologies.

    Good knowledge of infrastructure configuration management tools like Ansible, Chef, Puppet, and Terraform.

    Experience in using observability and tracing-related tools like InfluxDB, Prometheus, and Elastic stack.

    Ways to stand out from the crowd:
    Demonstrated experience in having SRE mindset, customer-first approach, and focus on customer satisfaction and passion for ensuring customer success.
    Experience with Git, code review, pipelines, and CI/CD.

    Interest in crafting, analyzing, and fixing large-scale distributed systems. Strong debugging skills with a systematic problem-solving approach to identify complex problems.

    Thrive in collaborative environments and enjoy working with various teams. Experience in using or running large private and public cloud systems based on Kubernetes, OpenStack, and Docker. Flexible in adapting to different working styles.

    NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and talented people on the planet working for us. If you're creative and autonomous, we want to hear from you
    The base salary range is 148,000 USD - 276,000 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

    You will also be eligible for equity and benefits .

    NVIDIA accepts applications on an ongoing basis.
    NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer.

    As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

    Similar Jobs (5)

    Senior Manager - Storage Production Engineering and SRE

    locations

    US, CA, Santa Clara

    time type

    Full time

    posted on

    Posted 4 Days Ago

    Senior HPC Storage Engineer

    locations

    3 Locations

    time type

    Full time

    posted on

    Posted 3 Days Ago

    Senior DevOps Engineer - Accelerated Computing

    locations

    5 Locations

    time type

    Full time

    posted on

    Posted 16 Days Ago
    NVIDIA pioneered accelerated computing to tackle challenges no one else can solve. Our work in AI and the metaverse is transforming the world's largest industries and profoundly impacting society.

    #J-18808-Ljbffr

  • Omega Solutions

    SRE Engineer

    3 weeks ago


    Omega Solutions Santa Clara, United States

    SRE Engineer - 01 Positions · St Louis, MO (Onsite from day 1) · Client · Required Skills: · •Bachelor's Degree in Computer Science, Computer Systems, Information Technology or related. Equivalent experience is acceptable. · •Experience with web applications and distributed ...

  • Omega Solutions

    SRE Engineer

    4 days ago


    Omega Solutions Santa Clara, United States

    SRE Engineer - 01 Positions · St Louis, MO (Onsite from day 1) · Client · Required Skills: · •Bachelor's Degree in Computer Science, Computer Systems, Information Technology or related. Equivalent experience is acceptable. · •Experience with web applications and distributed sy ...

  • Diverse Lynx

    SRE Engineer

    4 days ago


    Diverse Lynx Sunnyvale, United States

    Job Description: · SRE is a critical and visible role, central to running a multi-tiered cloud infrastructure, applications and workloads across public, private and hybrid cloud environments. SRE's are required to have in-depth knowledge of Cloud technologies. SRE's collaborate w ...

  • Diverse Lynx

    SRE Engineer

    3 weeks ago


    Diverse Lynx Sunnyvale, United States

    Job Description: · SRE is a critical and visible role, central to running a multi-tiered cloud infrastructure, applications and workloads across public, private and hybrid cloud environments. SRE's are required to have in-depth knowledge of Cloud technologies. SRE's collaborate ...

  • SOFTPATH TECHNOLOGIES

    SRE Engineer

    6 days ago


    SOFTPATH TECHNOLOGIES Sunnyvale, United States

    Job Description · Job Description · Hi All, · Greetings · PFB urgent W2 contract requirement and revert with your updated resume and current location ASAP to · Call me @ · SRE Engineer · Sunnyvale, CA Onsite -5 Days a Week · Face to Face Onsite Interview . · Local or Nea ...

  • Diverse Lynx

    SRE Engineer

    3 weeks ago


    Diverse Lynx Mountain View, United States

    Job Role: SRE Engineer · Location: Remote · Mountain View, California · Job Description · As a Site Reliability Engineer, be responsible to ensure optimal performance and up-time of Products critical security engineering services and infrastructure. Analyze system performan ...


  • Canonical - Jobs San Jose, CA, United States

    Job Description This role is an opportunity for a hands-on, but literally hands-off, technologist with a passion for Linux to build a career with Canonical and drive the success with those leveraging Ubuntu and open source products. If you have experience of IT operations automat ...

  • Wipro

    SRE Engineer

    3 weeks ago


    Wipro Mountain View, United States

    Wipro Limited (NYSE: WIT, BSE: 507685, NSE: WIPRO) is a leading technology services and consulting company focused on building innovative solutions that address clients most complex digital transformation needs. We leverage our holistic portfolio of capabilities in consulting, de ...

  • Diverse Lynx

    SRE Engineer

    4 days ago


    Diverse Lynx Mountain View, United States

    Job Role: SRE Engineer · Location: Remote · Mountain View, California · Job Description · As a Site Reliability Engineer, be responsible to ensure optimal performance and up-time of Products critical security engineering services and infrastructure. Analyze system performance ...

  • Wipro

    SRE Engineer

    3 weeks ago


    Wipro Mountain View, CA, United States

    Wipro Limited (NYSE: WIT, BSE: 507685, NSE: WIPRO) is a leading technology services and consulting company focused on building innovative solutions that address clients' most complex digital transformation needs. We leverage our holistic portfolio of capabilities in consulting, d ...

  • Redolent Infotech Pvt. Ltd.

    SRE DevOps Engineer

    3 weeks ago


    Redolent Infotech Pvt. Ltd. Sunnyvale, United States

    One of our direct client is urgently looking for a · SRE DevOps Engineer · @ Sunnyvale, CA · TITLE:SRE DevOps Engineer · LOCATION: Sunnyvale, CA · Duration: 6 to 12+ Months · Rate: DOE · Key Skills: · Splunk, Grafana, SRE, Cloud, DevOps, Azure, Docker, KuberNetes, Java (B ...

  • Info Way Solutions

    SRE Engineer

    1 week ago


    Info Way Solutions Fremont, United States

    Hi Professionals, · Hope you are doing good · This is Sangeetha from Info Way Solutions, LLC We have job opening for SRE Engineer and the detailed Job description is given below: · Kindly check the JD and share your views · Job Role: SRE Engineer · Location: Austin,Tx · Mandato ...

  • Info Way Solutions

    SRE Engineer

    4 days ago


    Info Way Solutions Fremont, United States

    Hi Professionals, · Hope you are doing good · This is · Sangeetha · from Info Way Solutions, LLC We have job opening for SRE Engineer and the detailed Job description is given below: · Kindly check the JD and share your views · Job Role: SRE Engineer · Location: Austin,Tx · ...


  • Wise Skulls llc Sunnyvale, United States

    Job Description · Job DescriptionTitle: SRE Engineering Program Manager · Location: Sunnyvale, CA (Hybrid) · Duration: 6+ months · Implementation Partner: Infosys · End Client: To be disclosed · JD: Primarily role of SRE EPM:SRE Team maintains Hadoop cluster and provides 24X7 sup ...


  • Wise Skulls llc Sunnyvale, United States

    Job Description · Job Description Title: SRE Engineering Program Manager · Location: Sunnyvale, CA (Hybrid) · Duration: 6+ months · Implementation Partner: Infosys · End Client: To be disclosed · JD: · Primarily role of SRE EPM: · SRE Team maintains Hadoop cluster and provides ...

  • Randstad

    sre engineer

    4 weeks ago


    Randstad San Leandro, United States

    sre engineer. · + san leandro , california · + posted april 15, 2024 · **job details** · summary · + $50 - $53 per hour · + contract · + bachelor degree · + category computer and mathematical occupations · + reference1048931 · job details · job summary: · In this contingent resou ...

  • Info Way Solutions

    SRE Engineer

    3 weeks ago


    Info Way Solutions Fremont, United States

    Experienced SRE consultant with hands on experience on setting up SRE Platform, defining SLI/SLO, Hands on experience on multiple Observability, Monitoring, and logging tools to help build MVP. · Required Skills Sets: · •Hands on with Java/Python, NoSQL, CI/CD, Azure/GCP · •Hand ...

  • Saxon Global

    SRE Engineer

    4 days ago


    Saxon Global Pleasanton, United States

    Key Responsibilities include, but are not limited to: · •Design, implement, test, and maintain container platforms like Pivotal Cloud Foundry and Azure Kubernetes Services. · •Ensure accessibility, security, reliability, availability, and performance of infrastructure. · •Admi ...

  • Randstad

    Sre engineer

    3 weeks ago


    Randstad San Leandro, United States

    job summary: · In this contingent resource assignment, you may: Consult on or participate in moderately complex initiatives and deliverables within Software Engineering and contribute to large-scale planning related to Software Engineering deliverables. Review and analyze moderat ...

  • Randstad

    Sre engineer

    3 weeks ago


    Randstad San Leandro, United States

    job summary: · In this contingent resource assignment, you may: Consult on or participate in moderately complex initiatives and deliverables within Software Engineering and contribute to large-scale planning related to Software Engineering deliverables. Review and analyze moderat ...