Data Center Facility Telemetry & Controls Engineer

Emploi pas sur LinkedIn

🕒 il y a 1 mois

🏄 California – Distant

info

💵 $185 000 - $290 000 / an

⏰ Temps Plein

🟠 Senior

🔴 Expert

👷🏻‍♀️ Ingénieur

🦅 Parrain de Visa H1B

info

🗣️🇺🇸🇬🇧 Anglais requis

Postuler Maintenant
Trouver des Emplois à Distance Similaires

📊 Vérifiez votre score de CV pour ce poste

Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

Logo of Lambda

Lambda

51 - 200 employés

💼 Conseil

🏥 Santé

📦 Logistique

💰 €39 700 000 Venture Round en 2022-11

Consulting • Healthcare • Logistics

Lambda est une entreprise d'informatique en nuage qui fournit des instances GPU à la demande et des clusters adaptés pour l'entraînement et l'inférence en IA. Elle propose une variété de produits GPU, tels que des instances GPU en nuage à la demande facturées à la minute, des clusters GPU privés à grande échelle et des serveurs PCIe avec des GPU NVIDIA Tensor Core personnalisables. Lambda est connue pour son cloud dédié aux développeurs IA, permettant aux développeurs d'IA de lancer des instances GPU en mettant l'accent sur le matériel le plus récent de NVIDIA. L'entreprise propose également des produits de station de travail configurés avec des GPU NVIDIA conçus pour l'apprentissage profond et d'autres applications d'IA.

Description

• Architect and manage BMS integration across colocation and Lambda-owned facilities, covering chillers, CRAHs, CDUs (Coolant Distribution Units), cooling towers, UPS systems, PDUs, and automatic transfer switches. • Define standards for BMS point lists, naming conventions, control sequences, and integration protocols (BACnet, Modbus, SNMP, OPC-UA, RESTful APIs). • Oversee commissioning and acceptance testing of new BMS deployments and CDU/TCS loop integrations for next-generation liquid-cooled GPU rack systems. • Collaborate with colocation partners (Equinix, Digital Realty, and others) to ensure telemetry data flows from provider BMS/EPMS into Lambda's monitoring stack. • Own the DCIM platform strategy and roadmap — evaluating, selecting, and implementing tooling for asset management, capacity planning, environmental monitoring, and power chain visibility. • Develop and maintain real-time dashboards for PUE, thermal performance, stranded capacity, and cooling system efficiency across all Lambda sites. • Build and maintain telemetry pipelines ingesting data from BMS, PDUs, in-rack sensors, CDUs, and network devices into centralized monitoring and alerting platforms (e.g., Prometheus, Grafana, InfluxDB, or equivalent). • Define alarm thresholds and escalation workflows for critical facility events including high coolant temperatures, CDU inlet/outlet anomalies, leak detection, and power exceedances. • Develop control strategies and setpoint frameworks for TCS (Thermal Control System) loops supporting direct liquid cooling at densities of 220–380 kW per rack. • Evaluate and qualify CDU vendors on controls integration capabilities, telemetry exposure, and remote management interfaces. • Define and enforce operational procedures for CDU commissioning, setpoint changes, loop pressure management, and fluid quality monitoring. • Support design and construction coordination for liquid cooling infrastructure in new data center buildouts, ensuring BMS and controls readiness at Day 1. • Establish and maintain facility event management processes, including on-call response protocols for facility telemetry anomalies. • Lead root cause analysis for facility system failures and implement corrective actions to prevent recurrence. • Partner with the data center operations team to maintain and refine emergency response runbooks tied to BMS alerts and automated controls. • Drive continuous improvement in MTTR for facility-related events through better telemetry coverage and automated remediation. • Manage BMS integrators, DCIM vendors, and control subcontractors - from RFP through design, installation, commissioning, and ongoing support. • Serve as the primary technical interface with colocation providers on all BMS/EPMS integration topics. • Collaborate with Lambda's infrastructure engineering, construction, and procurement teams to align controls requirements with facility buildout timelines. • Support due diligence and technical evaluation for new colocation sites and modular data center deployments from a telemetry and controls readiness perspective.

🎯 Exigences

• 7+ years of experience in data center infrastructure engineering, with at least 4 years focused on BMS, DCIM, or controls systems in a hyperscale, colocation, or AI/HPC environment. • Hands-on experience designing and integrating BMS for mission-critical facilities including UPS, PDU, CRAH/CRAC, chiller plant, cooling tower, and liquid cooling (CDU/in-row) systems. • Strong working knowledge of industrial control protocols: BACnet IP/MS-TP, Modbus TCP/RTU, SNMP, DNP3, and modern API-based integrations. • Demonstrated experience with DCIM platforms (Nlyte, Sunbird, Vertiv TRELLIS, or equivalent) including deployment, configuration, and ongoing administration. • Experience with real-time telemetry stacks (Prometheus, InfluxDB, Grafana, or similar) applied to infrastructure monitoring use cases. • Strong understanding of data center power and cooling systems, including PUE optimization, thermal management, and redundancy architectures (2N, N+1).

🏖️ Avantages

• Opportunity to shape the telemetry and controls architecture for one of the fastest-growing AI infrastructure platforms in the industry. • Work with cutting-edge GPU infrastructure at rack densities at the frontier of what the industry has deployed. • Collaborative environment with experienced infrastructure, construction, and vendor teams across a rapidly scaling global portfolio. • Competitive compensation including salary, equity, and comprehensive benefits. • Flexibility in work location with hybrid/remote options depending on facility portfolio needs.

Postuler Maintenant

Emplois Similaires

🕒 il y a 1 mois

Lumos

51 - 200

🤝 À but non lucratif

🤲 Charité

🌍 Impact social

OSP Project Engineer at Lumos responsible for planning and preparing construction drawings for fiber internet infrastructure. Designing for optimal use of communications facilities in a rapidly growing company.

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Greenhouse Software

501 - 1000

💼 Conseil

📣 Marketing

📦 Logistique

Revenue Intelligence Engineer optimizing internal applications in the Revenue Operations team at Greenhouse. Collaborating with tech and business leaders to enhance productivity across teams.

🇺🇸 États-Unis – Télétravail

💵 $128 300 - $180 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

👷🏻‍♀️ Ingénieur

🦅 Parrain de Visa H1B

info

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Ensono

1001 - 5000

💼 Conseil

Senior IAM Engineer overseeing the operational maintenance and expansion of ForgeRock IAM platform. Ensuring high availability and optimal performance while developing custom scripts and configurations.

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Rune Technologies

11 - 50

📦 Logistique

🏭 Fabrication

🎖️ Défense

Forward Deployed Engineer at Rune Technologies developing software solutions for military logistics. Collaborating with teams to deliver high-stakes projects and field-test systems.

🇺🇸 États-Unis – Télétravail

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

👷🏻‍♀️ Ingénieur

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Tern

11 - 50

💼 Conseil

📦 Logistique

📣 Marketing

Implementation Tooling Engineer at Tern, enhancing agency migrations through bulk operations. Focused on backend engineering and data pipeline optimization.

🇺🇸 États-Unis – Télétravail

💵 $140 000 - $170 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

👷🏻‍♀️ Ingénieur

🗣️🇺🇸🇬🇧 Anglais requis