AI Solution Architect

Job not on LinkedIn

🔥 42 minutes ago

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

🔴 Lead

💻 Solutions Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Uvation

Uvation

11 - 50 employees

💼 Consulting

🏥 Healthcare

🤖 Artificial Intelligence

Consulting • Healthcare • Artificial Intelligence

Uvation is a comprehensive technology services company providing innovative solutions across the IT infrastructure space. They specialize in artificial intelligence and cutting-edge cybersecurity solutions, offering products powered by Dell and Nvidia AI. Uvation delivers various managed services, including IT operations, security operations, and network operations, optimizing performance and enhancing security for businesses. Their services extend into public clouds with partnerships involving major platforms like Oracle Cloud, Google Cloud, and Amazon AWS. Uvation also operates a marketplace offering competitive pricing on a wide range of hardware and software products and provides a rewards program to incentivize customer engagement.

📋 Description

• Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness. • Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing and high-performance computing. • Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch and GPU resource allocation. • Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN and leaf-spine architectures. • Design AI storage and data architectures using object storage and parallel file systems such as Ceph and WEKA. • Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms and enterprise AI frameworks. • Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery and operational resilience. • Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations and implementation roadmaps. • Lead technical evaluations, proof-of-concepts, vendor assessments and architecture review boards. • Collaborate with infrastructure, network, security, storage, cloud, data, application and operations teams. • Define performance, availability, scalability, security and cost objectives and validate architecture against measurable acceptance criteria. • Provide technical leadership during deployment, migration, integration, troubleshooting and production transition. • Produce AI Factory reference architectures, solution blueprints, HLDs, LLDs, architecture diagrams, capacity/performance/scalability models, technology evaluations, security architecture inputs, BOMs, sizing, migration strategies, implementation roadmaps, operational readiness checklists, runbooks and acceptance criteria.

🎯 Requirements

• 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure. • Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures. • Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms. • Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred. • NVIDIA certifications or equivalent GPU/AI infrastructure credentials preferred. • AWS Solutions Architect / Azure Solutions Architect certification preferred. • TOGAF or equivalent enterprise architecture certification preferred. • CCNP/CCIE or equivalent networking certification preferred. • CISSP or equivalent security certification preferred. • Kubernetes certifications such as CKA/CKAD preferred. • Red Hat / Linux certifications preferred. • AI/ML architecture expertise, including NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystems. • PyTorch, TensorFlow and JAX, with operational understanding of training and inference workloads. • GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement. • LLM, generative AI, RAG, fine-tuning, model serving and inference architecture. • NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture. • DGX/HGX/OEM GPU server architecture and lifecycle management. • AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy. • 100/200/400/800G Ethernet, InfiniBand, RoCEv2, RDMA, Netris, BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN and QoS. • NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies. • GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting. • Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines. • Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies. • Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle. • GPUDirect Storage and storage/network performance optimization. • Kubernetes, GPU Operator, container runtimes, Kubernetes GPU scheduling and HPC or equivalent workload schedulers. • Model serving/inference platforms, MLOps platform architecture, API gateways, service discovery, secrets management and platform integration. • AWS and/or Azure AI infrastructure and security services. • Hybrid cloud connectivity, IAM, private networking, cloud storage, workload placement, cloud cost optimization, capacity planning and FinOps. • Zero Trust, network segmentation, IAM/RBAC, PAM, workload identity, GPU/DPU/container/Kubernetes/firmware/supply-chain security. • Encryption at rest/in transit, secrets management, audit logging, compliance controls, data/model protection, tenant isolation and secure model access. • Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry. • Monitoring across GPU, CPU, memory, network, storage, power and thermal domains. • High availability, backup/restore, disaster recovery, business continuity, failure-domain design, performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.

Apply Now

Similar Jobs

🕒 2 days ago

Grafana Labs

501 - 1000

🏢 Enterprise

☁️ SaaS

🤖 Artificial Intelligence

Senior Solutions Engineer partnering with Sales to sell Grafana Labs’ open-source observability platform. Delivering technical demos, evaluations, and customer enablement across India.

Grafana

Open Source

🕒 2 days ago

Incorta

201 - 500

💼 Consulting

🏭 Manufacturing

📦 Logistics

Solution Engineer implementing Incorta’s hyper-converged analytics platform for complex customer data projects. Designing solutions, teaching users, and supporting implementation delivery.

Amazon Redshift

AWS

Azure

Cloud

Linux

MySQL

Oracle

Python

Scala

Spark

SQL

🕒 6 days ago

MoneyGram

1001 - 5000

💳 Fintech

₿ Crypto

👥 B2C

Senior Integration Solutions Manager leading MoneyGram’s API, payment-rail, and partner integrations in India. Managing solution architecture, onboarding, compliance, testing, and go-live delivery.

Java

Python

.NET

🕒 September 11

Ookla

201 - 500

📡 Telecommunications

🏢 Enterprise

Customer Solutions Manager supporting Ookla’s connectivity intelligence customers. Troubleshooting technical issues, driving adoption, and using AI, analytics, and data products to grow strategic accounts.

SQL

Tableau

🕒 September 3

Grantek

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Technical Solutions Consultant delivering MES/MOM, Industrial DataOps, and digital transformation solutions for Grantek. Integrating manufacturing systems and leading technical implementations for global consumer brands.

ERP

Python

SQL