Software Engineer III – AI/ML Platform Operations

Job not on LinkedIn

🕒 July 2

🇺🇸 United States – Remote

💵 $105.3k - $140.6k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of AAA

AAA

5001 - 10000 employees

Founded 1902

🚘 Automotive

🛡️ Insurance

📦 Logistics

Automotive • Insurance • Logistics

AAA is a federation of affiliated automobile clubs, offering various travel-related services and products primarily for its members. It provides emergency roadside assistance, travel planning tools, and member discounts both domestically and internationally. Each club operates independently within its own geographical area, serving its local members while adhering to AAA's overarching standards. AAA also sells travel products like maps and guides, offering resources to assist with international travel. The organization facilitates a network where membership benefits, such as global discounts and emergency services, are extended to international members traveling in the United States and vice versa.

📋 Description

• Provide technical leadership for AI/ML platforms including Palantir, AWS Bedrock, Amazon SageMaker, and related cloud-native technologies • Ensure reliability, scalability, performance, security, and operational readiness for production AI workloads • Support deployment, monitoring, maintenance, and lifecycle management of AI/ML solutions and Generative AI services • Establish operational standards, support models, service-level objectives, and platform governance practices • Design and implement automation, monitoring, observability, and operational tooling • Develop and maintain dashboards, health metrics, alerts, logging frameworks, and operational runbooks • Enhance CI/CD pipelines, deployment automation, infrastructure-as-code, and model release processes • Implement MLOps, model monitoring, model lifecycle management, and AI operational governance practices • Serve as a senior escalation point for complex production issues involving AI platforms, machine learning workloads, cloud infrastructure, and data integrations • Lead root cause analysis and corrective and preventive actions • Resolve performance, availability, deployment, and integration issues across AI and data ecosystems • Partner with engineering and business teams to restore service and minimize operational risk • Mentor and provide technical guidance to engineers and platform teams • Influence platform strategy, architecture decisions, operational processes, and technology adoption • Collaborate across Data Engineering, Data Science, Architecture, Infrastructure, Security, and Product teams • Identify opportunities to improve operational efficiency, governance, security, and developer experience • Champion automation, observability, DevOps, Site Reliability Engineering, and AIOps practices • Travel as needed for divisional, team, and other in-person meetings

🎯 Requirements

• 3+ years of progressive experience in software engineering, platform engineering, cloud operations, MLOps, DevOps, or related technical disciplines • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience • Experience supporting production cloud-based applications and services in AWS environments • Strong software engineering and automation experience using Python, Java, JavaScript/TypeScript, or Node.js • Experience with CI/CD, build, integration, and deployment tools such as Jenkins, Maven, GitHub Actions, or equivalent • Experience with cloud-native compute, storage, networking, databases, and serverless architectures • Experience building and maintaining monitoring, observability, and alerting solutions • Strong troubleshooting, incident response, and root cause analysis skills • Excellent communication, collaboration, and technical leadership capabilities • Experience with AI/ML platforms such as Palantir Foundry, Amazon SageMaker, AWS Bedrock, Databricks, or similar ecosystems • Experience supporting Generative AI applications, LLM-based solutions, prompt orchestration frameworks, and RAG architectures • Knowledge of MLOps practices including model deployment, monitoring, governance, experimentation, and lifecycle management • Experience with Datadog, Splunk, Grafana, Prometheus, CloudWatch, or OpenTelemetry • Familiarity with AI governance, responsible AI principles, model risk management, and operational controls • Relevant cloud, AI/ML, DevOps, or platform engineering certifications • Must have authorization to work indefinitely in the United States • CSAA does not provide visa sponsorship; applicants must not require immigration support now or in the future • Ability to travel as needed for role, including divisional, team, and other in-person meetings

🏖️ Benefits

• Total compensation package • Annual bonus eligibility for most roles • 401(k) with a company match • Career growth and development opportunities • Leaders and mentors supporting long-term success • Remote-first Flexible Workplace • Home-Flex roles, working primarily from home • Flexibility to work from various locations including CSAA offices • Employee resource group participation • Volunteering opportunities • Company-wide annual discretionary bonus through the Annual Incentive Plan (AIP), up to 8% of eligible pay • Mentoring, leadership programs, tuition support, and hands-on experiences • Inclusive and welcoming workplace • Reasonable accommodations for qualified applicants and employees with disabilities or other limitations

Apply Now

Similar Jobs

🕒 July 2

Capital One

10,000+ employees

🏦 Banking

💳 Fintech

💸 Finance

Lead Full Stack Engineer at Capital One Shopping for diverse technology projects. Drive major transformation with cloud-based solutions while mentoring development teams.

🕒 July 2

Penn Mutual

1001 - 5000

Senior Software Engineer leading complex software systems design for financial applications. Collaborating with teams and mentoring engineers in software development and emerging technologies.

🕒 July 1

Signet Jewelers

10,000+ employees

🛒 Retail

🛍️ eCommerce

👗 Fashion

Lead Full Stack Developer at Signet Jewelers, responsible for driving design and development of enterprise applications. Collaborate with stakeholders to deliver scalable and secure solutions.

🕒 July 1

Signet Jewelers

10,000+ employees

🛒 Retail

🛍️ eCommerce

👗 Fashion

Lead Full Stack Developer at Signet Jewelers, guiding enterprise application development with a focus on .NET and AWS. Provide technical leadership and mentor development teams.

🕒 July 1

VetsEZ

201 - 500

🏥 Healthcare

💼 Consulting

📦 Logistics

Applied AI Product Engineer developing AI-driven healthcare technology platforms. Collaborating with teams to enhance patient engagement and operational efficiency remotely.