Enterprise Architect - AI
- Hiring Organisation
- Jobleads-UK
- Location
- Greater London, England, United Kingdom
record leading technical delivery (not just advisory) on enterprise-scale AI or HPC infrastructure programmes. AI Hardware & Data Center Infrastructure GPU/accelerator architectures: NVIDIA/AMD, including multi-node scale-out design. Accelerator interconnects: NVLink, NVSwitch High-performance networking: InfiniBand and RoCEv2 fabric design, 400G/800G Ethernet, rail … frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent multi-node training frameworks. Inference & serving: NVIDIA Triton, vLLM, TensorRT-LLM, or equivalent high-throughput serving platforms. MLOps/LLMOps: Kubeflow, MLflow, and at least one hyperscaler ML platform (SageMaker, Azure ...