Senior AI Platform Engineer – Infrastructure Services

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $132k - $182k / year

⏰ Full Time

🟠 Senior

👷 Infrastructure Engineer

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of SentinelOne

SentinelOne

1001 - 5000 employees

Founded 2013

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

SentinelOne is a leader in autonomous cybersecurity, known for its innovative use of AI across endpoint, cloud, and identity protection solutions. It is recognized by Gartner as a leader in the Magic Quadrant for Endpoint Protection Platforms for four consecutive years. SentinelOne's Singularity platform integrates enterprise security, offering features like AI-powered threat detection, endpoint and cloud security, vulnerability management, and threat intelligence. The company supports various industries by delivering real-time protection and operational efficiency while leveraging AI for advanced threat hunting and log analytics. With a strong focus on reducing risk and enhancing security performance, SentinelOne caters to enterprises worldwide with secure, scalable solutions.

📋 Description

• Take ownership of the AI Gateway infrastructure built on Kong AI Gateway, including authentication, routing, rate limiting, and monitoring of organization-wide AI coding assistant traffic • Architect, harden, and scale the Kong AI Gateway deployment (Konnect Hybrid on KCP/EKS) • Lead root-cause analysis, remediation, incident response, monitoring, and alerting for gateway reliability issues • Design solutions across CI/CD, GitOps, Kubernetes deployment tooling, artifact management, GitHub Enterprise, and GitHub Actions runner infrastructure • Evaluate and roll out AI developer tooling, including structured pilots and adoption efforts • Define AI infrastructure architecture and standards, review designs, and mentor engineers • Partner with security, DevEx, and product engineering teams to translate needs into platform capabilities • Deploy and operate self-hosted/open-weight model serving infrastructure, including GPU capacity planning, autoscaling, and cost/performance tuning • Build LLMOps practices covering model versioning, evaluation, safe rollout, vector stores, and embedding pipelines • Build observability for token usage, latency, and spending across API-based and self-hosted models

🎯 Requirements

• 5 or more years of experience in platform, infrastructure, or DevOps engineering • Track record of owning production systems end-to-end • Hands-on experience with API gateway technologies such as Kong, Envoy, or Apigee • Strong Kubernetes and GitOps experience, including ArgoCD or comparable tools • Experience operating across dev, gov, and prod environments • Jenkins pipeline design and administration experience • Build infrastructure and runner/agent fleet management experience • Experience with Artifactory, Xray, or similar artifact and package management systems • GitHub Enterprise administration experience • Working knowledge of Terraform and AWS/EKS • Experience deploying and operating self-hosted LLM inference stacks and GPU-backed infrastructure • Knowledge of Kubernetes GPU scheduling and autoscaling • Familiarity with LLMOps practices, model versioning, evaluation harnesses, and usage/cost observability • Track record of setting technical direction, driving cross-team initiatives, and mentoring engineers • Clear, proactive communication skills for explaining infrastructure trade-offs to technical and non-technical stakeholders • Experience with AI-assisted developer tooling at scale is preferred • Familiarity with Okta/OIDC and enterprise authentication patterns is preferred • Experience with LinearB or similar engineering productivity metrics tooling and Qodo or similar AI code review tooling is preferred • Experience with vector databases and RAG pipelines is preferred • Exposure to LoRA/QLoRA or similar model fine-tuning pipelines is preferred

🏖️ Benefits

• Restricted Stock Units (RSUs) • Employee Stock Purchase Plan (ESPP) • Flexible time off • Paid company holidays and paid sick time • Gender-neutral parental leave • Grandparent leave • Medical, dental, and vision coverage • 401(k) retirement plan with company match • Life and disability insurance • Health and dependent care FSA • Voluntary benefits (hospital, accident, critical illness) • Employee Assistance Program (EAP) • ARAG pre-paid legal • Nationwide pet insurance • Cancer Care program • Global business travel medical insurance • Home office allowance • Mobile phone reimbursement • Wellness coach • Wellness/gym reimbursement • Fertility coverage • Adoption & surrogacy reimbursement

Apply Now

Similar Jobs

🔥 12 hours ago

Hinshaw & Culbertson LLP

501 - 1000

⚖️ Legal

🛡️ Insurance

🏥 Healthcare

Senior infrastructure engineer designing and supporting servers, virtualization, storage, security, and Azure platforms for national law firm Hinshaw & Culbertson. Leading technical projects, escalations, monitoring, and business continuity initiatives.

🔥 18 hours ago

Unity

5001 - 10000

🏭 Manufacturing

💼 Consulting

📣 Marketing

Senior ML Engineer building reliable, low-latency online model inference infrastructure for Unity’s game engine and 3D creation platform. Optimizing serving, observability, experimentation, and deployment at scale.

🔥 22 hours ago

Affirm

1001 - 5000

💳 Fintech

👥 B2C

🛍️ eCommerce

Senior backend engineer building scalable batch infrastructure for Affirm, a buy-now-pay-later company. Designing distributed systems and workflow platforms supporting ML, product, and financial engineering.

🔥 23 hours ago

Netflix

10,000+ employees

📱 Media

👥 B2C

Full stack engineer building Netflix’s messaging and communications platforms for content production. Architecting scalable frontend and backend systems that power global entertainment workflows.

🕒 Yesterday

Mercola

201 - 500

🛍️ eCommerce

👥 B2C

🧘 Wellness

Systems Engineer modernizing Mercola’s production infrastructure across development, cloud, networking, identity, databases, and automation. Improving reliability through CI/CD, scripting, deployments, troubleshooting, and operational documentation.