AI Inference Engineer Jobs Dubai 2026

Fuse Energy — a forward-thinking renewable energy startup that has raised $210M from world-class investors including Multicoin, Balderton, Accel, and strategic angels including Nico Rosberg and co-founders behind Solana, Meta, Revolut, Spotify, and Uber — is hiring a Founding AI Inference Engineer based in Dubai. This is a rare, ground-floor opportunity to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. You will own the inference serving layer — request routing, batching, scheduling, autoscaling, and model optimisation — at a company sitting at the explosive intersection of renewable energy and AI compute infrastructure.

About Fuse Energy — Terawatt Mission Meets AI Compute

Mission: Deliver a terawatt of renewable energy — fast, using first-principles thinking and cutting-edge technology

New Frontier: Building GPU/CUDA performance layer AND inference serving layer simultaneously — from scratch

Market: Data centres are the fastest-growing electricity demand source — Fuse pairs real power with real compute

Position: Few companies can deliver both power and high-performance compute — Fuse is one of them

This Role: Founding engineer owning the inference serving layer — reporting directly to CTO

Funding, Investors & Benefits

Total Raised: $210M from Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group

Strategic Angels: Nico Rosberg · Solana Co-Founder · GPs behind Meta, Revolut, Spotify, Uber

Compensation: Competitive salary + equity sign-on bonus + biannual bonus scheme

Perks: Fully expensed tech to match your needs + breakfast and dinner allowance for office-based staff

Reporting: Direct to CTO — high ownership, no layers, real strategic impact from day one

Position Overview

The Founding AI Inference Engineer at Fuse Energy is one of the most compelling and consequential AI engineering roles available in Dubai in 2026. You will define the company’s entire inference serving strategy from first principles — designing and building the serving stack that handles request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive AI inference workloads across Fuse’s rapidly growing data centre infrastructure.

This is not a role where you inherit an established system and maintain it. You are the founding engineer who builds the system — making the core software architecture decisions about serving frameworks, model-level optimisation strategies, and performance standards that will define Fuse’s inference serving capability for years to come. You will work hand-in-hand with GPU and CUDA engineers, translate SLA commitments into technical specifications, and set the benchmarks and tooling standards for an entire new engineering function.

Why This Role Matters: The explosion in AI inference demand is one of the defining technology stories of the 2020s — and data centres are at the heart of it. Fuse Energy’s rare ability to pair genuine power delivery with high-performance compute creates an inference serving opportunity that very few companies in the world can match. As the Founding AI Inference Engineer, you are building the technical foundation that turns that advantage into the best inference offering in the market. This is not background infrastructure work. This is the product.

Key Responsibilities

Inference Serving Strategy & Architecture

  • Define Fuse’s inference serving strategy and architecture from first principles — making the foundational decisions about how AI models are served, scaled, and optimised across Fuse’s compute infrastructure with full technical ownership and accountability for outcomes
  • Make the core software architecture decisions on serving frameworks and orchestration — evaluating and selecting from tools such as vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalent frameworks based on Fuse’s specific performance, scalability, and reliability requirements
  • Set the standards, tooling, and benchmarks that the inference serving function will operate on as it scales — establishing engineering foundations that will support a growing team and increasing inference workload volume over time
  • Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans — converting commercial SLA promises into engineering architecture decisions that can actually be delivered and measured

Serving Stack Design & Engineering

  • Design and build the complete inference serving stack — covering request routing, dynamic batching, scheduling algorithms, and autoscaling logic for high-throughput, latency-sensitive inference workloads running across Fuse’s GPU cluster infrastructure
  • Own model-level optimisation strategy for serving — deciding where and how to apply techniques including quantisation, knowledge distillation, speculative decoding, and KV-cache management to improve throughput, reduce cost per token, and meet latency targets
  • Partner directly with CUDA and GPU engineering teams to ensure that custom kernels and hardware-level performance optimisations are cleanly and effectively integrated into the serving layer — maintaining clear ownership boundaries and efficient collaboration protocols
  • Reason clearly about the full inference path from incoming request to served response across a large GPU cluster — applying systems engineering thinking to identify and address bottlenecks in routing, batching, memory management, and response delivery

Performance Ownership & Reliability

  • Act as direct technical owner of inference performance and reliability — owning the metrics, the incident response process, and the continuous improvement cycle that keeps Fuse’s inference serving performing against committed SLAs and customer expectations
  • Design and implement autoscaling and capacity planning strategies for large-scale inference workloads — ensuring the serving infrastructure scales smoothly to meet demand spikes without compromising latency or cost efficiency
  • Build robust monitoring, benchmarking, and alerting infrastructure for inference serving — giving the team and leadership clear, real-time visibility of serving performance, capacity utilisation, and reliability metrics across the full system
  • Identify and address performance degradation risks proactively — before they impact serving SLAs or customer experience — through systematic load testing, capacity modelling, and architectural review processes

Founding Engineer Responsibilities

  • Shape the culture, engineering process, and technical standards of Fuse’s inference serving function from the ground up — as a founding engineer, your approach to code quality, system design, documentation, and technical decision-making sets the template for everyone who joins after you
  • Work directly with the CTO — contributing strategic technical input on inference serving architecture, compute investment, and the product roadmap for Fuse’s AI compute offering — with the influence and visibility that only a founding engineering role provides
  • Build and maintain clear technical documentation of the inference serving architecture, design decisions, and operational runbooks — ensuring that critical institutional knowledge is captured and accessible as the team grows

Requirements & Qualifications

Experience Requirements

  • 4+ years of experience building or operating large-scale inference serving systems — or equivalent strong project or industry experience that demonstrates the depth of inference serving knowledge this founding role requires
  • Deep, hands-on experience with inference serving frameworks and optimisation techniques — specifically batching strategies, KV-cache management, quantisation, and speculative decoding in production or near-production inference environments
  • A proven track record of making high-stakes architecture calls and owning the outcome — demonstrating the technical confidence and accountability that Fuse needs from the engineer who will define the entire inference serving function

Technical Skills

  • Strong systems thinking — able to reason about the full path from incoming inference request to served response across a large GPU cluster, identifying bottlenecks, trade-offs, and optimisation opportunities at every layer
  • Comfortable working directly with GPU/CUDA engineers to integrate low-level hardware performance work into a serving system — able to bridge the abstraction boundary between kernel engineering and serving architecture effectively
  • Comfort operating without a playbook — this is a founding role shaping a new function around architecture that is still early-stage, requiring the ability to make confident decisions under ambiguity and own those decisions through execution

Nice To Have

  • Experience with Triton or custom ML inference/training frameworks — indicating deeper capability with kernel-level or framework-level inference optimisation
  • Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system — providing directly transferable experience with the scale of challenges Fuse is building toward
  • Familiarity with Kubernetes or Slurm for cluster orchestration — supporting effective management of GPU cluster workloads alongside the inference serving stack
  • Interest or experience in energy markets, grid systems, or sustainability-focused compute — aligning with Fuse’s core mission and contributing to the team’s broader energy-meets-AI vision

Why This Founding AI Inference Role at Fuse Energy Is Exceptional in 2026

The convergence of renewable energy and AI compute infrastructure is one of the most significant technological and economic opportunities of the decade — and Fuse Energy is positioned at the exact intersection of both. With $210M raised from some of the world’s most respected technology investors, a clear terawatt energy delivery mission, and a rapidly growing data centre business driven by AI inference demand, Fuse is not a speculative bet. It is a well-capitalised, technically serious company building something that genuinely matters.

For an AI Inference Engineer with the right experience and the founding mindset, this role offers an extraordinary combination: direct CTO reporting, full architectural ownership, competitive salary plus equity, biannual bonus, expensed tech, and the mission-level credibility of being the person who built Fuse’s inference serving function from scratch. Based in Dubai with its zero income tax environment and rapidly growing AI ecosystem, this is one of the most compelling founding engineering opportunities available anywhere in the world in 2026.

Who Should Apply?

  • Senior AI Inference Engineers: With 4+ years of hands-on experience building or operating large-scale inference serving systems using vLLM, TensorRT-LLM, SGLang, Triton, or equivalent frameworks at production scale
  • ML Systems Engineers from Hyperscalers or AI Labs: With frontier-scale inference experience at companies like Google, Meta, OpenAI, Anthropic, Mistral, or equivalent — ready to own inference serving architecture at a well-funded, high-growth startup
  • GPU/CUDA-Fluent ML Engineers: Who understand how to bridge hardware-level performance work into serving-layer architecture — and who want to be the person who defines that integration from the ground up
  • Founding Engineers Seeking High-Impact Roles: Who are excited by the ownership, influence, and long-term equity upside of building a new engineering function at a $210M-funded company sitting at the intersection of renewable energy and AI
  • Dubai-Based or Relocating AI Engineers: Ready to join one of the most well-capitalised and technically ambitious AI-meets-energy startups in the world, reporting directly to the CTO and shaping what inference serving looks like at Fuse from day one

Recently Opening Job👇

AI Operations Engineer Jobs Dubai UAE 2026

JavaScript Developer Jobs UAE 2026

Leave a Comment

Select Your Degree:
Please select an option.
Select Your Experience:
Please select an option.
Select Currently Your Location:
Please select an option.
Please wait...
7
Aap ka agla page 7 second mein khulega...