HireFT
Browse JobsHow it worksPricingAboutSuccess Stories
    Back to jobs
    UV

    Uvation

    Technology

    AI Solution Architect

    Any, IndiaOn-SiteFull-time10+ yrs experiencePosted 1w ago
    View all jobs

    Job description

    Job Overview

    We are seeking an experienced AI Solution Architect to design and lead end-to-end enterprise AI Factory and GPU infrastructure solutions spanning compute, high-performance networking, storage, Kubernetes, cloud, and AI/ML platforms. The role requires strong expertise in NVIDIA GPU technologies, AI workloads, scalable infrastructure architecture, security, observability, performance engineering, and capacity planning.

    Key Responsibilities

    • Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
    • Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing, and high-performance computing.
    • Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch, and GPU resource allocation.
    • Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.
    • Design AI storage and data architectures using object storage, parallel file systems like Ceph, WEKA, , or equivalent platforms.
    • Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.
    • Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.
    • Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations, and implementation roadmaps.
    • Lead technical evaluations, proof-of-concepts, vendor assessments, and architecture review boards.
    • Collaborate with infrastructure, network, security, storage, cloud, data, application, and operations teams.
    • Define performance, availability, scalability, security, and cost objectives and validate architecture against measurable acceptance criteria.
    • Provide technical leadership during deployment, migration, integration, troubleshooting, and production transition.

    • Required Technical Skills

    AI / ML Architecture

    • NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystem.

    • PyTorch, TensorFlow, JAX and operational understanding of training and inference workloads.

    • GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.

    • LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.

    GPU & AI Factory Infrastructure

    • NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; familiarity with next-generation systems.

    • NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.

    • DGX/HGX/OEM GPU server architecture and lifecycle management.

    • AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.

    High-Performance Networking

    • 100/200/400/800G Ethernet, InfiniBand, RoCEv2 and RDMA, Netris

    • NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.

    • BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, QoS and congestion management.

    • GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.

    AI Storage & Data Architecture

    • Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.

    • Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.

    • Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.

    • GPUDirect Storage and storage/network performance optimization.

    AI Platform & Orchestration

    • Kubernetes, GPU Operator, container runtimes and Kubernetes GPU scheduling.

    • HPC or other equivalent workload schedulers.

    • Model serving/inference platforms and MLOps platform architecture.

    • API gateways, service discovery, secrets management and platform integration.

    Cloud & Hybrid Architecture

    • AWS and/or Azure AI infrastructure and security services.

    • Hybrid cloud connectivity, IAM, private networking, cloud storage and workload placement.

    • Cloud cost optimization, capacity planning and FinOps considerations for GPU workloads.

    Security & Governance

    • Zero Trust, network segmentation, IAM/RBAC, PAM and workload identity.

    • GPU, DPU, container, Kubernetes, firmware and supply-chain security.

    • Encryption at rest/in transit, secrets management, audit logging and compliance controls.

    • AI-specific risks including data/model protection, tenant isolation and secure model access.

    Observability & Reliability

    • Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.

    • Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.

    • High availability, backup/restore, disaster recovery, business continuity and failure-domain design.

    • Performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.

    Architecture Deliverables

    • AI Factory reference architecture and solution blueprints

    • High-Level Design (HLD) and Low-Level Design (LLD)

    • Network, compute, GPU and storage architecture diagrams

    • Capacity, performance and scalability models

    • Technology evaluation and vendor comparison documents

    • Security architecture and threat-model inputs

    • Bill of Materials (BOM) and infrastructure sizing

    • Migration/deployment strategy and implementation roadmap

    • Operational readiness checklist, runbooks and acceptance criteria

    Experience & Qualifications

    • 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.

    • Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.

    • Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.

    • Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.

    Preferred Certifications

    • NVIDIA certifications or equivalent GPU/AI infrastructure credentials

    • AWS Solutions Architect / Azure Solutions Architect

    • TOGAF or equivalent enterprise architecture certification

    • CCNP/CCIE or equivalent networking certification

    • CISSP or equivalent security certification

    • Kubernetes certifications such as CKA/CKAD

    • Red Hat / Linux certifications

    Job details are sourced from the employer's original posting.

    Open job posting
    UV

    About the company

    Uvation

    Uvation is a technology company specializing in AI and machine learning solutions.

    View all Uvation jobs
    Industry
    Technology
    Open roles
    93

    Interested in this role?

    Apply with HireFT

    Free to start — no card required.

    Your fit

    How well do you match?

    Sign in to see how your résumé lines up with this role.