Full Time

Lead Applied AI Site Reliability Engineer II

Deloitte
Nashville, TN
$150,000 - $220,000* / year

Job Description

About this kind of role

What a Lead Applied AI Site Reliability Engineer II usually does. The employer's own details are in the listing.

Typically does: This role focuses on ensuring the reliability, performance, and scalability of machine learning systems and infrastructure. Responsibilities include designing, building, and maintaining tools and processes for monitoring, incident response, and automation. The engineer will collaborate with machine learning engineers and other stakeholders to troubleshoot issues and improve system resilience. They also contribute to the development of best practices for site reliability engineering within an applied AI context.

Tools and skills: Proficiency in cloud platforms (e.g., AWS, Azure, GCP), scripting languages (e.g., Python), containerization technologies (e.g., Docker, Kubernetes), monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack), and experience with infrastructure-as-code is generally expected.

Good fit for: Individuals with a strong background in both software engineering and site reliability, coupled with a passion for machine learning and a desire to build robust, scalable systems, typically thrive in this position.

View Similar Jobs

Matches Jobs

Similar jobs which you may be interested in. Typically using your existing skillset.