Gateworth Group is supporting a high-growth technology business building advanced accelerator hardware and large-scale AI infrastructure, seeking a Network CCL Engineer in Dubai who can operate deep in the performance layer. This role focuses on collective communication, accelerator runtime, heterogeneous device communication, and distributed AI workloads.
About the Role – Collective Communication for AI Accelerators
Core Focus: Collective communication primitives for accelerator-based systems
Technical Scope: Multi-device, rack-level inference, and heterogeneous accelerator/GPU communication
Key Libraries: NCCL, RCCL, oneCCL, MPI, UCX, or libfabric
Transport Mechanisms: PCIe, RDMA-based networking, custom fabrics, and device-to-device communication
Collaboration: Hardware, runtime, compiler, and distributed systems teams
Career Growth & Next-Gen AI Infrastructure Impact
Strategic Location: Dubai – a growing hub for advanced AI hardware and infrastructure development
Deep Technical Ownership: Build the communication foundations for next-generation AI platforms
Cutting-Edge Exposure: Work on LLM workload optimization including tensor parallelism, MoE, and KV-cache movement
Career Growth: Build rare, specialized expertise at the intersection of hardware and distributed AI systems
Position Overview
This Network CCL Engineer role develops collective communication primitives for accelerator-based systems, building communication support for multi-device and rack-level inference environments. You will design heterogeneous communication paths across accelerator and GPU devices, create runtime components for device discovery, rank management, and synchronization, integrate transport mechanisms across PCIe, RDMA, and custom fabrics, collaborate with hardware, runtime, and distributed systems teams, and diagnose and optimize communication bottlenecks including bandwidth, latency, and congestion.
Why This Role Matters: As Network CCL Engineer, you build the low-level communication foundations that determine how efficiently next-generation AI platforms scale across multi-chip and multi-rack environments, apply deep expertise in collective communication operations like AllReduce and AllToAll to real accelerator hardware, work at the critical intersection of hardware, runtime, and distributed systems engineering, optimize communication bottlenecks that directly impact training and inference performance at scale, and gain rare, highly specialized experience in a niche that sits at the core of the AI infrastructure boom.
Key Responsibilities
Collective Communication Development
- Develop collective communication primitives for accelerator-based systems
- Build communication support for multi-device and rack-level inference environments
- Design heterogeneous communication paths across accelerator and GPU devices
Runtime Component Development
- Create runtime components for device discovery, rank management, and topology awareness
- Build queues, streams, events, and synchronization mechanisms
Transport Integration
- Integrate transport mechanisms across PCIe, RDMA-based networking, and custom fabrics
- Support device-to-device communication across heterogeneous hardware
Cross-Team Collaboration & Optimization
- Collaborate with hardware, runtime, compiler, and distributed systems teams on scalable communication solutions
- Diagnose and optimize communication bottlenecks, including bandwidth, latency, congestion, and synchronization overhead
Qualifications & Requirements
Experience Requirements
- 5+ years in distributed systems, runtime engineering, networking, or accelerator-level development
Programming Skills
- Strong C/C++ and Python programming experience
Technical Knowledge
- Solid understanding of collective communication operations (AllReduce, ReduceScatter, AllGather, AllToAll, Broadcast, Send/Recv, Barrier)
- Familiarity with libraries such as NCCL, RCCL, oneCCL, MPI, UCX, or libfabric
- Understanding of GPU/NPU memory behavior, memory registration, and peer-to-peer transfers
- Experience with runtime concepts including queues, streams, events, and command submission
- Understanding of hardware-software interaction: DMA engines, device memory, interrupts, doorbells, and fences
- Knowledge of high-performance networking technologies (InfiniBand, RoCE)
Bonus Experience
- Experience with distributed training frameworks or custom communication backends
- Exposure to communication optimization for LLM workloads (tensor parallelism, MoE, KV-cache movement)
- Experience integrating communication layers with modern distributed inference frameworks
About This Opportunity via Gateworth Group
Gateworth Group is supporting a high-growth technology business building advanced accelerator hardware and large-scale AI infrastructure. This Network CCL Engineer role offers a progressive and challenging opportunity to work deep in the performance layer of next-generation AI platforms, collaborating closely with hardware, runtime, and distributed systems teams on some of the most technically demanding engineering problems in the AI infrastructure space.
Career Excellence: Build the collective communication foundations powering next-generation AI accelerator platforms in Dubai.
Who Should Apply?
- Collective Communication Engineers: With NCCL, RCCL, or MPI experience in production systems
- Distributed Systems Engineers: Strong in C/C++ and Python with runtime engineering background
- Accelerator/GPU Software Engineers: Familiar with GPU/NPU memory behavior and peer-to-peer transfers
- High-Performance Networking Specialists: Experienced with InfiniBand, RoCE, and RDMA
- LLM Infrastructure Engineers: Interested in tensor parallelism, MoE, and KV-cache optimization
Recently Opening Job👇
Gen AI Solution Architect Intern Jobs Dubai UAE 2026
Specialist Cloud Security AI Security Jobs Abu Dhabi UAE 2026
