AI & HPC Infrastructure Engineer

Job not on LinkedIn

🔥 2 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of FirstPrinciples

FirstPrinciples

11 - 50 employees

Founded 2024

🤖 Artificial Intelligence

🔬 Science

☁️ SaaS

Artificial Intelligence • Science • SaaS

FirstPrinciples is a research company building AI systems for discovery in fundamental science. It develops domain-specialized models and tools (branded Theo: Theo Collaborator, Theo Conjecture, and Theo, the AI Physicist) to assist the scientific process—hypothesis generation, symbolic reasoning, tool integration, validation loops, and reproducible research objects. The company emphasizes transparency, stewardship of knowledge as a public good, and alignment with the scientific community. Technical claims include multiple fine-tuned models across physics domains, a model family with 120B+ parameters trained on a curated corpus of 3M+ scientific papers, and internal experiments in areas like quantum information.

📋 Description

• Design, deploy, and operate Kubernetes infrastructure for AI inference, research, and engineering workloads • Set up and manage GPU and HPC-style compute environments, including scheduling, utilization, job management, and node-level troubleshooting • Work with systems such as Kubernetes, Slurm or similar schedulers, container runtimes, GPU drivers & libraries (ie; CUDA), storage systems, and observability tools • Build and manage Linux-based compute environments, including provisioning, networking, storage, monitoring, access control, and lifecycle management • Help architect bare metal, cloud, and hybrid infrastructure across AWS, GCP, Azure, or equivalent platforms • Own the reliability and operational health of infrastructure systems, including monitoring, alerting, incident response, capacity planning, and performance tuning • Improve deployment workflows, automation, configuration management, secrets management, and infrastructure-as-code practices • Partner with ML engineers, researchers, and software engineers to understand workload requirements and translate them into practical infrastructure designs • Evaluate tradeoffs between managed cloud services, self-managed Kubernetes, HPC schedulers, bare metal deployments, and multi-cloud architectures • Build tooling, documentation, runbooks, and operational practices that help the team move quickly without making infrastructure fragile or opaque • Balance speed and robustness, knowing when to prototype quickly and when to harden systems for long-term use

🎯 Requirements

• Strong infrastructure builder with experience operating production, research, cloud, or high-performance compute systems • Deeply comfortable with Linux administration, including debugging networking, storage, system services, permissions, performance issues, and node-level failures • Experienced with Kubernetes in real environments, including cluster operations, deployments, networking, observability, scaling, and troubleshooting • Comfortable working with cloud infrastructure on AWS, GCP, Azure, or equivalent platforms • Familiar with infrastructure automation and configuration tools such as Terraform, Ansible, Helm, ArgoCD, GitOps workflows, or similar systems • Experienced with GPU-heavy, compute-heavy, or HPC-style workloads, especially in environments involving AI, ML, research computing, or scientific workloads • Able to work across bare metal and cloud environments, and interested in the practical tradeoffs between the two • Comfortable reasoning about resource scheduling, cluster utilization, autoscaling, storage, networking, and observability for distributed workloads • Practical and ownership-oriented; you can take ambiguous infrastructure needs and turn them into working systems • Comfortable collaborating across disciplines, especially with researchers and engineers who may not think in infrastructure terms • Able to operate independently as a senior or strong intermediate contributor, while knowing when to bring others into important technical decisions • Motivated by building foundational systems that make ambitious technical and scientific work possible

🏖️ Benefits

• The opportunity to work on foundational problems at the intersection of AI and physics • A high-trust, low-bureaucracy environment with real ownership • Remote-first work with flexibility in how you structure your day • Exposure to cutting-edge ideas across AI, scientific discovery, infrastructure, and emerging technologies • A culture that values curiosity, depth of thinking, and first-principles reasoning • The chance to shape the compute and inference infrastructure behind advanced AI systems for scientific discovery

Apply Now

Similar Jobs

🕒 4 days ago

ClickHouse

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Senior Cloud Data Infrastructure Engineer at ClickHouse responsible for developing cloud-native database solutions. Collaborating with teams to enhance cloud infrastructure and scalability.

AWS

Azure

Cloud

Distributed Systems

EC2

Google Cloud Platform

Java

Kafka

Kubernetes

Spark

Go

🕒 July 16

Fullscript

201 - 500

🏥 Healthcare

📦 Logistics

🏭 Manufacturing

Senior Backend Engineer for infrastructure team at Fullscript. Responsible for core application's health, scalability, and developer experience.

🇨🇦 Canada – Remote

💵 $125k - $165k / year

💰 $240M Private Equity Round on 2021-11

⏰ Full Time

🟠 Senior

👷 Infrastructure Engineer

Ruby

Ruby on Rails

🕒 June 30

AbacusNext

201 - 500

⚖️ Legal

💼 Consulting

☁️ SaaS

Senior Infrastructure Engineer managing IT systems and architecture across various components. Leading projects and collaborating within Agile/SCRUM teams while ensuring system performance and stability.

Ansible

Azure

Firewalls

Terraform

VMware

🕒 June 25

Mechanical Orchard

11 - 50

💼 Consulting

📦 Logistics

📣 Marketing

Senior Infrastructure Software Engineer at Mechanical Orchard building modernization platform for critical business applications. Collaborating with clients and teams to deploy IMOGEN in cloud environments.

Cloud

🕒 June 20

decircle

1 - 10

📣 Marketing

📦 Logistics

💼 Consulting

Senior Infrastructure Engineer collaborating with teams to design scalable infrastructure solutions for TRM Labs' AI-powered platforms. Involves software development, monitoring, and compliance with standards.

Airflow

AWS

BigQuery

Google Cloud Platform

Kubernetes

Linux

React

Terraform

Unix