Senior Solutions Architect, First Time Deployment Validation

🕒 vor 2 Monaten

🏄 California, Texas, +2 weitere Bundesländer – Remote

infoinfo

💵 $148.000 - $235.750 / Jahr

⏰ Vollzeit

🟠 Senior

💻 Lösungsingenieur

🦅 H1B-Visum-Sponsor

infoinfo

👻 Geisterscore 18%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of NVIDIA

NVIDIA

10.000+ Mitarbeiter

Gegründet 1993

🏥 Gesundheitswesen

🏭 Fertigung

🤖 Künstliche Intelligenz

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA ist ein führendes Technologieunternehmen mit Spezialisierung auf beschleunigtes Computing und Künstliche Intelligenz (AI). NVIDIA treibt Fortschritte bei Grafikprozessoren (GPUs), Cloud Computing, Rechenzentren und Virtual Reality voran und fokussiert dabei Branchen wie Gaming, Automotive, Gesundheitswesen und Robotik. Innovationen des Unternehmens wie NVIDIA Omniverse transformieren traditionelle digitale Prozesse, indem sie hochrealistische Simulationen und Rendering-Aufgaben ermöglichen. Die Anwendungen erstrecken sich über zahlreiche Branchen – von autonomen Fahrzeugen mit NVIDIA DRIVE über Gesundheitslösungen mit NVIDIA Clara bis hin zu AI-gestützten Analysen und Workflows.

Beschreibung

• Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters. • Ensure configurations align with guidelines for NCCL, collectives, and distributed training frameworks. • Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis. • Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform. • Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health. • Develop automation (Python, Shell) for running benchmarks, collecting results, and performing regression checks • Examine communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll. • Recommend changes to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency. • Work closely with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use. • Contribute to documentation, guidelines, and readiness collateral that support internal collaborators and customer-facing teams.

🎯 Anforderungen

• Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field. • More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings. • Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, with practical knowledge of NCCL. • Solid grasp of collective communication patterns, particularly AllReduce and AllToAll, and how they are applied in contemporary ML/LLM training. • Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or TensorFlow. • Proficiency with Python and Shell/Bash for scripting, automation, and tooling. • Experience with benchmarking (crafting, executing, and interpreting performance benchmarks). • Comfortable working with observability data (metrics, logs, dashboards) to troubleshoot and optimize complex distributed workloads. • Strong communication skills and the ability to work effectively with cross-functional teams.

🏖️ Vorteile

• equity • benefits

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 2 Monaten

Clutch

51 - 200

📣 Marketing

✈️ Reisen

💼 Beratung

Senior Solutions Engineer bridging sales and technical needs for WithClutch. Fostering client relationships and demonstrating product solutions in fintech.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 2 Monaten

S-Docs

51 - 200

☁️ SaaS

🏢 Unternehmen

🤝 B2B

Solution Architect designing, configuring, and delivering S-Docs solutions across onboarding, professional services, and custom implementations for clients. Collaborating with teams to ensure successful deployment and customer success.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

💻 Lösungsingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

AWS

Azure

Cloud

ETL

Informatica

Java

Python

SOAP

Visualforce

🕒 vor 2 Monaten

Cloudelligent

51 - 200

💼 Beratung

☁️ SaaS

AWS Solutions Architect at Cloudelligent responsible for cloud modernization, data engineering, and AI/ML integration. Collaborating with clients to create innovative cloud architectures using AWS technologies.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

🔴 Experte

💻 Lösungsingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 2 Monaten

Cloudelligent

51 - 200

💼 Beratung

☁️ SaaS

AWS Solutions Architect at Cloudelligent responsible for cloud modernization and AI/ML solutions. Collaborating with clients to design and implement scalable cloud solutions using AWS technologies.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

🔴 Experte

💻 Lösungsingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 2 Monaten

SHI International Corp.

5001 - 10000

💼 Beratung

📦 Logistik

🏭 Fertigung

Senior Solutions Architect at SHI driving new business in Microsoft AI solutions. Collaborating closely with technical teams in a presales context to develop integrated solutions.

🇺🇸 Vereinigte Staaten – Remote

💵 $175.000 - $235.000 / Jahr

⏰ Vollzeit

🟠 Senior

💻 Lösungsingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich