AI Infrastructure Architect

🔥 0 minutes ago

🇮🇱 Israel – Remote

⏰ Full Time

🟠 Senior

🔴 Lead

👷 Infrastructure Engineer

👻 Ghost score 15%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NeuReality

NeuReality

51 - 200 employees

🏭 Manufacturing

💼 Consulting

📦 Logistics

💰 Series A on 2024-03

Manufacturing • Consulting • Logistics

NeuReality is a Series A venture capital-backed AI startup that revolutionizes AI Inferencing in data centers worldwide. Founded by seasoned engineers, their AI-centric system architecture—from silicon to software—delivers 10x performance or 90% cost savings, making AI Inferencing faster, affordable, and energy-efficient. With over 60 employees across 3 countries, NeuReality launched its flagship NR1 AI Inference Solution in 2023. Their solutions empower Cloud Service Providers and Enterprise Customers in sectors such as financial services, healthcare, and life sciences to boost AI Inferencing performance for large-scale AI pipelines. NeuReality simplifies complex AI deployments, advancing the state of affordable and sustainable AI.

📋 Description

• Lead the software architecture and technical roadmap for NeuReality’s NR-Nexus • Write system specifications for the NR-Nexus product • Research AI infrastructure, SaaS platforms, model serving, and inference trends • Work with engineering to translate technical capabilities into product value • Work closely with engineering teams to optimize performance, scalability, and feature delivery • Define performance goals and lead profiling, benchmarking, and optimization efforts for GenAI and distributed AI workloads • Collaborate with customers, partners, and open-source communities to ensure ecosystem compatibility and adoption • Mentor software engineers and provide technical leadership

🎯 Requirements

• 7+ years of software engineering experience, including 3+ years in software architecture or technical leadership • Strong experience with Kubernetes-based platforms and cloud-native architecture • Deep understanding of Gen AI/LLM infrastructure and distributed workloads • Experience designing management software or SaaS platforms for production systems • Strong background in distributed systems, microservices, APIs, and automation • Hands-on experience with observability stacks, monitoring, logging, alerting, and SLA/SLO tracking • Experience with CI/CD, deployment automation, upgrades, and rollback mechanisms • Good understanding of security, authentication, authorization, and integration with customer data center environments • Deep understanding of GenAI / LLM inference infrastructure, including model serving, scaling, batching, latency, throughput, and resource utilization • Experience with production AI inference clusters using GPUs, AI accelerators, or other specialized compute infrastructure • Understanding of how distributed inference systems operate, including scheduling, load balancing, autoscaling, failover, and cluster-level observability • Experience with LLM serving frameworks such as vLLM, Triton Inference Server, TensorRT-LLM, or similar • Familiarity with GPU/accelerator orchestration, device plugins, resource scheduling, and cluster capacity planning • Familiarity with GPU communication technologies such as GPUDirect RDMA, NCCL, NVLink, or UALink • Experience optimizing communication for distributed AI/ML workloads • Knowledge of Prometheus, Grafana, OpenTelemetry, Helm, Argo CD, Istio, KServe, Kubeflow, or similar tools • Experience deploying software in on-prem, edge, private cloud, or hybrid environments

Apply Now