Model Serving Infrastructure
16 products

Fireworks Inference Engine
Fireworks
·Inference APIs+2
-

AI Foundation
DataRobot
·Model Deployment Platforms+2
-

Voice AI Cloud
AssemblyAI
·Inference APIs+2
-

Dedicated Inference
CoreWeave
·Model Serving Infrastructure+2
-

AI Object Storage
CoreWeave
·Model Serving Infrastructure+2
-

CPU Compute
CoreWeave
·GPU Cloud Platforms+2
-

Docker Model Runner
Docker
·Model Deployment Platforms+2
-

Akamai Inference Cloud
Akamai
·Inference APIs+2
-

Data-Lakehouse & ML-Ops
Polynom
·Model Deployment Platforms+2
-

Merge Gateway
Merge
·Model Serving Infrastructure+3
-

Model Vault
Cohere
·Model Deployment Platforms+2
-

Data Centers
Galaxy
·Model Serving Infrastructure+2
-

NVIDIA HGX Platform
NVIDIA
·Inference Acceleration+2
-

GPU Instances
Scaleway
·GPU Cloud Platforms+2
-

Generative APIs - Dedicated Deployment
Scaleway
·Inference APIs+2
-

Inference Endpoints
Hugging Face
·Inference APIs+2
-

Inference APIsMainInference AccelerationModel Serving Infrastructure
An optimized inference engine that serves the latest open models or custom-trained versions with industry-leading throughput and latency, offering serverless, on-demand, and reserved deployment options.
Features
- Serve open models or custom-trained versions with optimized performance
- Deploy models with serverless, on-demand, or reserved capacity options
- Optimize inference with custom kernels and memory management
- Support multi-region deployments and custom performance optimizations
- Utilize disaggregated prefill and decode for reduced latency
- Implement speculative decoding and advanced caching strategies
- Enable multi-node expert parallelism for large models
- Provide day-zero support for new model releases

Inference APIsMainInference AccelerationModel Serving Infrastructure
An optimized inference engine that serves the latest open models or custom-trained versions with industry-leading throughput and latency, offering serverless, on-demand, and reserved deployment options.
Features
- Serve open models or custom-trained versions with optimized performance
- Deploy models with serverless, on-demand, or reserved capacity options
- Optimize inference with custom kernels and memory management
- Support multi-region deployments and custom performance optimizations
- Utilize disaggregated prefill and decode for reduced latency
- Implement speculative decoding and advanced caching strategies
- Enable multi-node expert parallelism for large models
- Provide day-zero support for new model releases