logo
Alauda AI
  • Alauda AI
  • Deploy
  • Inference Gateway
Navigation
Overview
Introduction
Architecture
Quick Start
Release Notes
Plan
Architecture, Version and Components
Validated Models
Qwen3-32B
Qwen3.6-27B (W8A8)
DeepSeek-V4-Flash (W4A8)
DeepSeek-V4-Flash (W8A8)
Gemma-4-31B-it
MiniMax-M2.5 (W8A8)
Device Options
About Alauda Build of Hami
About Alauda Build of NVIDIA GPU Device Plugin
Glossary
Install
Pre-installation Configuration
Install Workbench
Configure Alauda AI Entry and Branding
Install Alauda AI
Tools Menu Configuration
Upgrade
Upgrade Alauda AI
Migrating to Knative Operator
Upgrade Workbench
Uninstall
Administer
Hardware Profile
Introduction
Hardware Profile Management
Create Hardware Profile using CLI
Schedule Workloads to Specific GPU Nodes
Creating CPU-Only and GPU-Accelerated Profiles
Multi-Tenant
Namespace Management
MLflow Workspaces and Access Control
Secure Profile
Install the Secure-Profile Dependencies
Enable the Secure Profile
Develop
Workbench
Introduction
Create Workbench
Kubeflow Notebooks
Use Kubeflow Notebooks
Use Kubeflow Volumes
Use Kubeflow Tensorboards
Connections
Introduction
Using Connections
Pipelines
Create a Data Science Pipelines Application
Use Kubeflow Pipelines
Run Kubeflow Pipelines from JupyterLab with Elyra
Use Reusable Kubeflow Pipeline Components
Kubeflow Pipelines Execution and Storage Behavior
Schedule Alauda DevOps Pipelines with Kueue
Distributed Workloads
CodeFlare SDK Tutorial
Run a Spark Application
Use Kubeflow Model Registry
Experiment Tracking
Using the MLflow Python SDK with Authentication and RBAC
Kubeflow Pipeline + MLflow Integration
Trace AI agents with MLflow
Agentic MLOps
Use Coding Agents with On-Premise Inference Services
Run MLOps with Coding Agents and On-Premise LLMs
Components
Kubeflow
Alauda support for Kubeflow
Install Kubeflow Operators
FAQ
Upgrade Kubeflow Operators
Data Science Pipelines
Data Science Pipelines Operator
Installation
KubeRay
Alauda Build of KubeRay Operator
Installation
Spark Operator
Alauda Build of Spark Operator
Installation
MLflow
MLflow
Installation
Label Studio
Label Studio
Install Label Studio
Quickstart
Main Features
Feast
Alauda Build of Feast
Install Feast
Quickstart
Train
Training Guides
Kubeflow Trainer Quick Start
Checkpointing and Resuming TrainJobs
Preemptible TrainJobs with Kueue, Checkpointing, and Inference Coexistence
GPU Slicing with Dynamic Resource Allocation (DRA)
Training Runtime Images
Fine-tuning LLMs with Training Hub
Daily Fine-Tuning Pipeline with MLflow Tracking and TrustyAI Evaluation
Fine-tuning LLMs using Workbench
Fine-tune and Pretrain LLMs on Ascend NPU
Quota & Scheduling
Setup RBAC
Configuring quotas
Using cohorts
Configuring fair sharing
Gang scheduling
Manage Ascend NPU quota with Kueue
Components
Kueue
Alauda Build of Kueue
Install Kueue
Volcano
Alauda support for Volcano
Install Volcano
JobSet
Alauda Build of JobSet
Install JobSet
Quickstart
Deploy
Model Management
Introduction
Model Repository
Upload Models Using Notebook
Model Storage
Share Models
Inference Service
Introduction
Managing Inference Services
Guides
Create Inference Service using CLI
Deploy Inference Services from the Kubeflow Dashboard
Extend Inference Runtimes
Using KServe Modelcar for Model Storage
Configure External Access for Inference Services
Configure Scaling for Inference Services
Set Up Autoscaling for Inference Services with KEDA
Scheduling Inference Services based on the CUDA version
Schedule Inference Services with Kueue
Enable Expert Parallel for vLLM Inference Services
Speculative Decoding for vLLM Inference Services
Troubleshooting
Experiencing Inference Service Timeouts with MLServer Runtime
Inference Service Fails to Enter Running State
Model as a Service (MaaS)
Inference Gateway
Authenticating Consumers
Configuring Token Quotas
Metering Token Usage
Routing to LLM Providers
Charging Back Token Usage
LLM Compressor
Introduction
LLM Compressor with Alauda AI
Components
KServe
Alauda Build of KServe
Install KServe
Envoy AI Gateway
Alauda Build of Envoy AI Gateway
Install Envoy AI Gateway
LeaderWorkerSet
Alauda Build of LeaderWorkerSet
Install LeaderWorkerSet
InferNex Bridge
Alauda Build of InferNex Bridge
Install InferNex Bridge
Build AI Applications
Quick Start with Core Profile
Demo with Secure Profile
Demo Advanced: AuthBridge on the Tool with Token Exchange
Kagenti with two MCP servers behind Envoy AI Gateway
Components
Dify
Dify
Install Dify
Main Features
Llama Stack
Alauda Build of Llama Stack
Install Llama Stack
Quickstart
Main Features
Kagenti
Kagenti Operator
Installation
Security Architecture
MCP Lifecycle Operator
Alauda Build of MCP Lifecycle Operator
Installation
Quickstart
Evaluate & Safety
Evaluate LLM
Evaluating RAG with Ragas
AI Guardrails for LLM safety
NeMo Guardrails
Components
TrustyAI
Alauda Build of TrustyAI
Install TrustyAI
Deploy TrustyAI Service
Monitor
Logging & Tracing
Introduction
Logging
Resource Monitoring
Introduction
Monitoring Metrics and Views
Add a Monitoring Dashboard
Monitoring pending workloads
Monitor Dashboard Stuck at Loading
Bias and Drift Monitoring
API Reference
Introduction
Kubernetes APIs
Inference Service APIs
ClusterServingRuntime [serving.kserve.io/v1alpha1]
InferenceService [serving.kserve.io/v1beta1]
Workbench APIs
Workspace Kind [kubeflow.org/v1beta1]
Workspace [kubeflow.org/v1beta1]
Manage APIs
AmlNamespace [manage.aml.dev/v1alpha1]
Operator APIs
AmlCluster [amlclusters.aml.dev/v1alpha1]

#Inference Gateway

Previous pageModel as a Service (MaaS)Next pageAuthenticating Consumers