Deploy

This chapter covers creating and scaling inference services, model management and storage, Model as a Service (MaaS), quota and metering at the inference gateway, and model compression.

The components this chapter relies on are listed under Components below.

Model Management

Inference Service

Model as a Service (MaaS)

Inference Gateway

LLM Compressor

Components