Model Serving Infrastructure

16 products

Fireworks Inference Engine
Fireworks
·Inference APIs+2
-
AI Foundation
DataRobot
·Model Deployment Platforms+2
-
Voice AI Cloud
AssemblyAI
·Inference APIs+2
-
Dedicated Inference
CoreWeave
·Model Serving Infrastructure+2
-
AI Object Storage
CoreWeave
·Model Serving Infrastructure+2
-
CPU Compute
CoreWeave
·GPU Cloud Platforms+2
-
Docker Model Runner
Docker
·Model Deployment Platforms+2
-
Akamai Inference Cloud
Akamai
·Inference APIs+2
-
Data-Lakehouse & ML-Ops
Polynom
·Model Deployment Platforms+2
-
Merge Gateway
Merge
·Model Serving Infrastructure+3
-
Model Vault
Cohere
·Model Deployment Platforms+2
-
Data Centers
Galaxy
·Model Serving Infrastructure+2
-
NVIDIA HGX Platform
NVIDIA
·Inference Acceleration+2
-
GPU Instances
Scaleway
·GPU Cloud Platforms+2
-
Generative APIs - Dedicated Deployment
Scaleway
·Inference APIs+2
-
Inference Endpoints
Hugging Face
·Inference APIs+2
-
Inference APIsMainInference AccelerationModel Serving Infrastructure

An optimized inference engine that serves the latest open models or custom-trained versions with industry-leading throughput and latency, offering serverless, on-demand, and reserved deployment options.

Features

  • Serve open models or custom-trained versions with optimized performance
  • Deploy models with serverless, on-demand, or reserved capacity options
  • Optimize inference with custom kernels and memory management
  • Support multi-region deployments and custom performance optimizations
  • Utilize disaggregated prefill and decode for reduced latency
  • Implement speculative decoding and advanced caching strategies
  • Enable multi-node expert parallelism for large models
  • Provide day-zero support for new model releases