# Gateway API Inference Extension Type of document: How-to guide Product: NGINX Gateway Fabric --- Learn how to use NGINX Gateway Fabric with the Gateway API Inference Extension to optimize traffic routing to self-hosting Generative AI Models on Kubernetes. ## Overview The [Gateway API Inference Extension](https://gateway-api-inference-extension.sigs.k8s.io/) is an official Kubernetes project that aims to provide optimized load-balancing for self-hosted Generative AI Models on Kubernetes. The project's goal is to improve and standardize routing to inference workloads across the ecosystem. Coupled with the provided llm-d Router, NGINX Gateway Fabric becomes an [Inference Gateway](https://gateway-api-inference-extension.sigs.k8s.io/#concepts-and-definitions). An Inference Gateway adds AI-specific traffic management features, such as model-aware routing, serving priority for models, and model rollouts. ## Set up Install the Gateway API Inference Extension CRDs: ```shell kubectl kustomize "https://github.com/nginx/nginx-gateway-fabric/config/crd/inference-extension/?ref=v" | kubectl apply -f - ``` To enable the Gateway API Inference Extension, [install](/ngf/install/) NGINX Gateway Fabric with these modifications: - Using Helm: set the `nginxGateway.gwAPIInferenceExtension.enable=true` Helm value. - Using Kubernetes manifests: set the `--gateway-api-inference-extension` flag in the nginx-gateway container argument, update the ClusterRole RBAC to add the `inferencepools`: ```yaml - apiGroups: - inference.networking.k8s.io resources: - inferencepools verbs: - get - list - watch - apiGroups: - inference.networking.k8s.io resources: - inferencepools/status verbs: - update ``` See this [example manifest](https://raw.githubusercontent.com/nginx/nginx-gateway-fabric/main/deploy/inference/deploy.yaml) for clarification. ## Deploy a sample model server The [vLLM simulator](https://github.com/llm-d/llm-d-inference-sim) model server does not use GPUs and is ideal for test/development environments. To deploy the vLLM simulator, run the following command: ```shell kubectl apply -f https://raw.githubusercontent.com/kubernetes-sigs/gateway-api-inference-extension/refs/tags/v/config/manifests/vllm/sim-deployment.yaml ``` ## Deploy the InferencePool and llm-d router The InferencePool is a Gateway API Inference Extension resource that represents a set of Inference-focused Pods. With InferencePool, you can configure a routing extension as well as inference-specific routing optimizations. For more information on this resource, refer to the Gateway API Inference Extension [InferencePool documentation](https://gateway-api-inference-extension.sigs.k8s.io/api-types/inferencepool/). Install an InferencePool named `vllm-qwen3-32b` that selects from endpoints with label `app: vllm-qwen3-32b` and listening on port 8000. The Helm install command automatically installs the llm-d router and InferencePool. NGINX queries the llm-d Router to find the pod endpoint that gets the traffic. The llm-d Router picks from the ready pods that the InferencePool `selector` field matches. For more information, see the README for the llm-d router's [Endpoint Picker](https://github.com/llm-d/llm-d-router/blob/main/pkg/epp/README.md). **Note:** The llm-d router is a third-party application written and provided by the llm-d project. Communication between NGINX and the llm-d router uses TLS with certificate verification disabled by default. NGINX Gateway Fabric is not responsible for any threats or risks associated with using this third-party llm-d router application. **Note:** For all chart values, see the [llm-d Router Helm charts](https://github.com/llm-d/llm-d-router/tree/main/config/charts). ```shell helm install vllm-qwen3-32b \ --set router.modelServers.matchLabels.app=vllm-qwen3-32b \ --version v \ oci://ghcr.io/llm-d/charts/llm-d-router-gateway ``` **Note:** For test environments, lower the CPU and memory requests and limits to reduce resource use: ```shell helm install vllm-qwen3-32b \ --set router.modelServers.matchLabels.app=vllm-qwen3-32b \ --version v \ --set router.epp.resources.requests.cpu=100m \ --set router.epp.resources.requests.memory=512Mi \ --set router.epp.resources.limits.memory=2Gi \ oci://ghcr.io/llm-d/charts/llm-d-router-gateway ``` Confirm that the llm-d router was deployed and is running: ```shell kubectl describe deployment vllm-qwen3-32b-epp ``` ## Deploy an Inference Gateway ```yaml kubectl apply -f - < ``` ## Deploy an HTTPRoute ```yaml kubectl apply -f - <