
11 - 50 Mitarbeiter
Gegründet 2022
🤖 Künstliche Intelligenz
📚 Bildung
🤝 Non-Profit
Artificial Intelligence • Education • Non-profit
Ashby Electrical Limited ist ein 2018 gegründeter Elektrofachbetrieb, der sich auf kosteneffiziente und nachhaltige Elektrolösungen für Wohn- und Gewerbeprojekte spezialisiert hat. Das Unternehmen steht für hochwertige Arbeit über den gesamten Projektlebenszyklus hinweg – von der Planung über die Installation und Prüfung bis zur Inbetriebnahme. Ashby Electrical setzt auf starke, von Vertrauen, Integrität und Professionalität geprägte Kundenbeziehungen und stellt die termingerechte und budgetkonforme Umsetzung sicher.
🕒 vor 18 Tagen
🌐 Vereinigte Staaten, Singapur – Remote
🏄 California – Remote
💵 $150.000 - $275.000 / Jahr
⏰ Vollzeit
🟠 Senior
🧑💻 Full-Stack-Entwickler
👻 Geisterscore 14%
🗣️🇺🇸🇬🇧 Englisch erforderlich
Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

11 - 50 Mitarbeiter
Gegründet 2022
🤖 Künstliche Intelligenz
📚 Bildung
🤝 Non-Profit
Artificial Intelligence • Education • Non-profit
Ashby Electrical Limited ist ein 2018 gegründeter Elektrofachbetrieb, der sich auf kosteneffiziente und nachhaltige Elektrolösungen für Wohn- und Gewerbeprojekte spezialisiert hat. Das Unternehmen steht für hochwertige Arbeit über den gesamten Projektlebenszyklus hinweg – von der Planung über die Installation und Prüfung bis zur Inbetriebnahme. Ashby Electrical setzt auf starke, von Vertrauen, Integrität und Professionalität geprägte Kundenbeziehungen und stellt die termingerechte und budgetkonforme Umsetzung sicher.
• Operate the Kubernetes GPU fleet day to day, including node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning • Own batch scheduling and multi-tenancy, including queues, quotas, priorities, preemption, gang scheduling, and fair share across research teams • Design and run storage under the fleet, including high-performance shared filesystems, object storage tiers, quotas, and backups • Keep multi-node training runs fault-tolerant by owning node health and automated draining, debugging NCCL and fabric problems, tracking stragglers and flaky GPUs, and building checkpoint and restart patterns • Harden the platform through identity and access, network policy, secrets, workload isolation, and sandboxing for AI agents • Bring new capacity online by acceptance-testing providers on fabric, NCCL, and storage throughput, holding them to SLAs, and integrating new clusters with infrastructure as code • Work directly with research teams on infrastructure problems and turn recurring issues into platform fixes • Share the on-call rotation, runbooks, and postmortems • Collaborate with researchers and engineers to keep large-scale experiments performant and fault-tolerant
• 3+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms • Owned at least one system from design through operation • Production Kubernetes experience for GPU workloads with a batch layer such as Slurm, Kueue, Volcano, or similar • Experience with quotas, priority and preemption, and node health • Experience owning infrastructure as code and observability for a production fleet • Experience with Terraform or Ansible • Experience deploying with Helm and ArgoCD • Experience monitoring with Prometheus or equivalent tools • Strong programming skills in at least one infrastructure language such as Python, Go, Rust, or C++ • Ability to write clearly for engineers, researchers, and providers • Additional relevant expertise in distributed training infrastructure, distributed storage, cluster security, scheduler internals, or multi-provider platforms is advantageous
• Health Insurance - 94% of Insurance premium paid by Organization commencing within 1 month after your start date • Retirement - 401(k) plan with up to 2% match • 25 days Paid Time Off per year, accrued weekly • Up to 10 days of paid sick leave per year • Paid Bereavement, Family, Medical and Pregnancy Disability Leave • WFH stipend and work computer provided for eligible employees • Catered lunches and dinners on workdays at the Berkeley office • Work-related travel and equipment expenses are covered • Visa sponsorship for in-person employees
Jetzt Bewerben🕒 vor 18 Tagen
Senior Software Engineer building scalable Workday Finance services for DTN, a global operational data and technology company. Designing secure APIs, distributed systems, and financial workflows.
🇺🇸 Vereinigte Staaten – Remote
💵 $109.500 - $145.500 / Jahr
💰 Grant im 2014-09
⏰ Vollzeit
🟠 Senior
🧑💻 Full-Stack-Entwickler
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 18 Tagen
Lead Technical Consultant designing and implementing ServiceNow solutions for client organizations. Leading delivery teams and driving projects to successful completion.
🇺🇸 Vereinigte Staaten – Remote
💰 Post-IPO Equity im 2022-12
⏰ Vollzeit
🟠 Senior
🧑💻 Full-Stack-Entwickler
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 18 Tagen
Senior Software Engineer building Lithic’s real-time card authorization and fraud-prevention systems. Developing scalable payment-risk tools, fraud controls, and distributed backend services.
🇺🇸 Vereinigte Staaten – Remote
💵 $160.000 - $200.000 / Jahr
⏰ Vollzeit
🟠 Senior
🧑💻 Full-Stack-Entwickler
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 18 Tagen
Senior engineer building Cribl Stream integrations with Splunk, Kafka, and cloud storage. Developing NodeJS and TypeScript software for enterprise telemetry infrastructure.
🇺🇸 Vereinigte Staaten – Remote
💵 $160.000 - $220.000 / Jahr
⏰ Vollzeit
🟠 Senior
🧑💻 Full-Stack-Entwickler
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 18 Tagen
Senior Software Engineer building AI-powered Go automation for Chainguard’s secure container image factory. Designing developer tooling, validation systems, and agentic pipelines for open-source supply chain security.
🗣️🇺🇸🇬🇧 Englisch erforderlich