Course
Give your AI a permanent memory of your business. A course for people who use ChatGPT or Claude daily.Compound Context
Together AI logo

Together AI

4.4

Together AI is a cloud platform that helps developers and AI researchers run, train, and fine-tune open-source AI models. It offers access to powerful NVIDIA GPUs and a wide selection of models, allowing you to build and deploy your AI applications quickly and affordably with flexible options.

About Together AI

Who It's For

This platform helps developers, AI researchers, and companies with AI teams. It's great for those needing powerful GPUs and flexible tools to build custom AI applications. It is not for non-technical users looking for ready-to-use AI.

What You Get

You access powerful NVIDIA GPUs and a library of over 200 open-source AI models for tasks like chat or images. The platform offers tools to run models, fine-tune them with your data, or train new ones. You also get flexible options, from serverless to dedicated hardware.

How It Works

You connect via an easy API. Select models, customize them with your data using fine-tuning, or run them to power your apps. The system is designed for speed and cost-effectiveness, helping you deploy AI projects quickly and affordably.

Stay in the loop

Weekly roundup of new AI agents. No spam, unsubscribe anytime.

Subscribe and get the free 2026 AI Agents Field Guide

Join 1,500+ AI builders · weekly, no spam

Features & Capabilities

⚙️ Core AI Platform Services

Serverless Inference API

Deploy and run inference on open-source AI models without managing underlying servers.

Fine-Tuning Platform

Train and improve high-quality, task-specific models using your proprietary data.

Model Library

Explore and build with a diverse collection of open-source and specialized AI models.

Pre-Training Capabilities

Securely and cost-effectively train custom AI models from the ground up.

⚡ High-Performance GPU Infrastructure

Instant GPU Clusters

Access self-service NVIDIA GPUs instantly for rapid model deployment and training.

Dedicated Endpoints

Deploy models on custom hardware with guaranteed capacity and expert support.

Global Data Center Network

Leverage GPU power distributed across 25+ data center locations worldwide.

Frontier Hardware Support

Utilize cutting-edge NVIDIA GPUs like GB200 NVL72 for optimal performance.

🛠️ AI Development & Agentic Tools

Code Sandbox

Create secure and isolated development environments for building AI applications.

Code Interpreter

Safely execute LLM-generated code to enhance AI agent capabilities.

LLM Selection Tool

Find the most suitable large language model for your specific application requirements.

Advanced Tool Calling

Integrate external tools and functions into AI agents for complex use cases.

📊 Performance & Cost Optimization

ATLAS Inference System

Accelerate LLM inference by up to 4x using runtime-learning techniques.

Batch Inference API

Process billions of tokens at a 50% lower cost for efficient large-scale inference.

Optimized Unit Economics

Achieve industry-leading performance-to-cost ratios for AI workloads.

Production Scale Reliability

Ensures stable and consistent performance for high-volume, production-grade applications.

Use Cases

Accelerating Production AI Inference at Scale

Organizations face challenges deploying generative AI models with high performance and cost-efficiency in production. Together AI's serverless and dedicated inference endpoints, powered by innovations like ATLAS, deliver unmatched price-performance, enabling reliable deployment of open-source models for high-throughput and streaming applications.

B2B SaaSFor: AI/ML Engineers

Customizing and Deploying Task-Specific LLMs

Generic open-source models often lack the specificity required for unique business challenges. Together AI's fine-tuning platform allows developers to train models with proprietary data, creating highly accurate, task-specific, and cost-effective AI solutions that are fully owned and easily integrated into production workflows.

Any Industry leveraging Specialized AIFor: Machine Learning Engineers

Building and Executing AI Agents with Code Capabilities

Developers creating sophisticated AI agents need robust environments for code generation and execution. Together AI's Code Sandbox and Code Interpreter, combined with high-performance GPU clusters and support for advanced tool calling, provide the necessary infrastructure to rapidly build, test, and scale AI-native applications.

Software DevelopmentFor: AI Developers

Provisioning Enterprise-Grade GPU Infrastructure for AI Workloads

Large enterprises require secure, reliable, and massively scalable GPU clusters for pre-training custom models and deploying complex AI solutions globally. Together AI offers "Frontier AI Factory" and Reserved Clusters with NVIDIA's latest hardware across global data centers, providing the dedicated capacity and expert support needed for critical enterprise AI initiatives.

Large EnterpriseFor: Enterprise Architects

Frequently asked questions

Tags

Specifications

Deployment
Cloud
API
Browser
Target Audience
Individual
Startup
Business
Enterprise
Complexity
Expert

Pricing

Llama 4 Maverick (Input)

Per one-time

$0.27
  • Llama 4 Maverick Input usage
  • Price per 1M tokens

Llama 4 Maverick (Output)

Per one-time

$0.85
  • Llama 4 Maverick Output usage
  • Price per 1M tokens

Llama 4 Scout (Input)

Per one-time

$0.18
  • Llama 4 Scout Input usage
  • Price per 1M tokens

Llama 4 Scout (Output)

Per one-time

$0.59
  • Llama 4 Scout Output usage
  • Price per 1M tokens

Llama 3.3 70B Instruct-Turbo (Input)

Per one-time

$0.88
  • Llama 3.3 70B Instruct-Turbo Input usage
  • Price per 1M tokens

Llama 3.3 70B Instruct-Turbo (Output)

Per one-time

$0.88
  • Llama 3.3 70B Instruct-Turbo Output usage
  • Price per 1M tokens

Llama 3.2 3B Instruct Turbo (Input)

Per one-time

$0.06
  • Llama 3.2 3B Instruct Turbo Input usage
  • Price per 1M tokens

Llama 3.2 3B Instruct Turbo (Output)

Per one-time

$0.06
  • Llama 3.2 3B Instruct Turbo Output usage
  • Price per 1M tokens

Llama 3.1 405B Instruct Turbo (Input)

Per one-time

$3.5
  • Llama 3.1 405B Instruct Turbo Input usage
  • Price per 1M tokens

Llama 3.1 405B Instruct Turbo (Output)

Per one-time

$3.5
  • Llama 3.1 405B Instruct Turbo Output usage
  • Price per 1M tokens

Llama 3.1 70B Instruct Turbo (Input)

Per one-time

$0.88
  • Llama 3.1 70B Instruct Turbo Input usage
  • Price per 1M tokens

Llama 3.1 70B Instruct Turbo (Output)

Per one-time

$0.88
  • Llama 3.1 70B Instruct Turbo Output usage
  • Price per 1M tokens

Llama 3.1 8B Instruct Turbo (Input)

Per one-time

$0.18
  • Llama 3.1 8B Instruct Turbo Input usage
  • Price per 1M tokens

Llama 3.1 8B Instruct Turbo (Output)

Per one-time

$0.18
  • Llama 3.1 8B Instruct Turbo Output usage
  • Price per 1M tokens

Llama 3 8B Instruct Lite (Input)

Per one-time

$0.1
  • Llama 3 8B Instruct Lite Input usage
  • Price per 1M tokens

Llama 3 8B Instruct Lite (Output)

Per one-time

$0.1
  • Llama 3 8B Instruct Lite Output usage
  • Price per 1M tokens

Llama 3 70B Instruct Reference (Input)

Per one-time

$0.88
  • Llama 3 70B Instruct Reference Input usage
  • Price per 1M tokens

Llama 3 70B Instruct Reference (Output)

Per one-time

$0.88
  • Llama 3 70B Instruct Reference Output usage
  • Price per 1M tokens

Llama 3 70B Instruct Turbo (Input)

Per one-time

$0.88
  • Llama 3 70B Instruct Turbo Input usage
  • Price per 1M tokens

Llama 3 70B Instruct Turbo (Output)

Per one-time

$0.88
  • Llama 3 70B Instruct Turbo Output usage
  • Price per 1M tokens

LLaMA-2 (Input)

Per one-time

$0.9
  • LLaMA-2 Input usage
  • Price per 1M tokens

LLaMA-2 (Output)

Per one-time

$0.9
  • LLaMA-2 Output usage
  • Price per 1M tokens

DeepSeek-R1 (Input)

Per one-time

$3
  • DeepSeek-R1 Input usage
  • Price per 1M tokens

DeepSeek-R1 (Output)

Per one-time

$7
  • DeepSeek-R1 Output usage
  • Price per 1M tokens

DeepSeek R1 Distilled Qwen 14B (Input)

Per one-time

$0.18
  • DeepSeek R1 Distilled Qwen 14B Input usage
  • Price per 1M tokens

DeepSeek R1 Distilled Qwen 14B (Output)

Per one-time

$0.18
  • DeepSeek R1 Distilled Qwen 14B Output usage
  • Price per 1M tokens

DeepSeek R1 Distilled Llama 70B (Input)

Per one-time

$2
  • DeepSeek R1 Distilled Llama 70B Input usage
  • Price per 1M tokens

DeepSeek R1 Distilled Llama 70B (Output)

Per one-time

$2
  • DeepSeek R1 Distilled Llama 70B Output usage
  • Price per 1M tokens

DeepSeek R1-0528-tput (Input)

Per one-time

$0.55
  • DeepSeek R1-0528-tput Input usage
  • Price per 1M tokens

DeepSeek R1-0528-tput (Output)

Per one-time

$2.19
  • DeepSeek R1-0528-tput Output usage
  • Price per 1M tokens

DeepSeek-V3-1 (Input)

Per one-time

$0.6
  • DeepSeek-V3-1 Input usage
  • Price per 1M tokens

DeepSeek-V3-1 (Output)

Per one-time

$1.7
  • DeepSeek-V3-1 Output usage
  • Price per 1M tokens

DeepSeek-V3 (Input)

Per one-time

$1.25
  • DeepSeek-V3 Input usage
  • Price per 1M tokens

DeepSeek-V3 (Output)

Per one-time

$1.25
  • DeepSeek-V3 Output usage
  • Price per 1M tokens

gpt-oss-120B (Input)

Per one-time

$0.15
  • gpt-oss-120B Input usage
  • Price per 1M tokens

gpt-oss-120B (Output)

Per one-time

$0.6
  • gpt-oss-120B Output usage
  • Price per 1M tokens

gpt-oss-20B (Input)

Per one-time

$0.05
  • gpt-oss-20B Input usage
  • Price per 1M tokens

gpt-oss-20B (Output)

Per one-time

$0.2
  • gpt-oss-20B Output usage
  • Price per 1M tokens

Qwen3-Coder 480B A35B Instruct (Input)

Per one-time

$2
  • Qwen3-Coder 480B A35B Instruct Input usage
  • Price per 1M tokens

Qwen3-Coder 480B A35B Instruct (Output)

Per one-time

$2
  • Qwen3-Coder 480B A35B Instruct Output usage
  • Price per 1M tokens

Qwen3 235B A22B Instruct 2507 FP8 (Input)

Per one-time

$0.2
  • Qwen3 235B A22B Instruct 2507 FP8 Input usage
  • Price per 1M tokens

Qwen3 235B A22B Instruct 2507 FP8 (Output)

Per one-time

$0.6
  • Qwen3 235B A22B Instruct 2507 FP8 Output usage
  • Price per 1M tokens

Qwen3 235B A22B Thinking 2507 FP8 (Input)

Per one-time

$0.65
  • Qwen3 235B A22B Thinking 2507 FP8 Input usage
  • Price per 1M tokens

Qwen3 235B A22B Thinking 2507 FP8 (Output)

Per one-time

$3
  • Qwen3 235B A22B Thinking 2507 FP8 Output usage
  • Price per 1M tokens

Qwen3 235B A22B FP8 Throughput (Input)

Per one-time

$0.2
  • Qwen3 235B A22B FP8 Throughput Input usage
  • Price per 1M tokens

Qwen3 235B A22B FP8 Throughput (Output)

Per one-time

$0.6
  • Qwen3 235B A22B FP8 Throughput Output usage
  • Price per 1M tokens

Qwen 2.5 72B (Input)

Per one-time

$1.2
  • Qwen 2.5 72B Input usage
  • Price per 1M tokens

Qwen 2.5 72B (Output)

Per one-time

$1.2
  • Qwen 2.5 72B Output usage
  • Price per 1M tokens

Qwen2.5-VL 72B Instruct (Input)

Per one-time

$1.95
  • Qwen2.5-VL 72B Instruct Input usage
  • Price per 1M tokens

Qwen2.5-VL 72B Instruct (Output)

Per one-time

$8
  • Qwen2.5-VL 72B Instruct Output usage
  • Price per 1M tokens

Qwen2.5 Coder 32B Instruct (Input)

Per one-time

$0.8
  • Qwen2.5 Coder 32B Instruct Input usage
  • Price per 1M tokens

Qwen2.5 Coder 32B Instruct (Output)

Per one-time

$0.8
  • Qwen2.5 Coder 32B Instruct Output usage
  • Price per 1M tokens

Qwen2.5 7B Instruct Turbo (Input)

Per one-time

$0.3
  • Qwen2.5 7B Instruct Turbo Input usage
  • Price per 1M tokens

Qwen2.5 7B Instruct Turbo (Output)

Per one-time

$0.3
  • Qwen2.5 7B Instruct Turbo Output usage
  • Price per 1M tokens

Qwen QwQ-32B (Input)

Per one-time

$1.2
  • Qwen QwQ-32B Input usage
  • Price per 1M tokens

Qwen QwQ-32B (Output)

Per one-time

$1.2
  • Qwen QwQ-32B Output usage
  • Price per 1M tokens

GLM-4.5-Air (Input)

Per one-time

$0.2
  • GLM-4.5-Air Input usage
  • Price per 1M tokens

GLM-4.5-Air (Output)

Per one-time

$1.1
  • GLM-4.5-Air Output usage
  • Price per 1M tokens

Kimi K2 Instruct (Input)

Per one-time

$1
  • Kimi K2 Instruct Input usage
  • Price per 1M tokens

Kimi K2 Instruct (Output)

Per one-time

$3
  • Kimi K2 Instruct Output usage
  • Price per 1M tokens

Kimi K2 Thinking (Input)

Per one-time

$1.2
  • Kimi K2 Thinking Input usage
  • Price per 1M tokens

Kimi K2 Thinking (Output)

Per one-time

$4
  • Kimi K2 Thinking Output usage
  • Price per 1M tokens

Kimi K2 0905 (Input)

Per one-time

$1
  • Kimi K2 0905 Input usage
  • Price per 1M tokens

Kimi K2 0905 (Output)

Per one-time

$3
  • Kimi K2 0905 Output usage
  • Price per 1M tokens

Mistral (7B) Instruct v0.2 (Input)

Per one-time

$0.2
  • Mistral (7B) Instruct v0.2 Input usage
  • Price per 1M tokens

Mistral (7B) Instruct v0.2 (Output)

Per one-time

$0.2
  • Mistral (7B) Instruct v0.2 Output usage
  • Price per 1M tokens

Mistral Instruct (Input)

Per one-time

$0.2
  • Mistral Instruct Input usage
  • Price per 1M tokens

Mistral Instruct (Output)

Per one-time

$0.2
  • Mistral Instruct Output usage
  • Price per 1M tokens

Mistral Small 3 (Input)

Per one-time

$0.8
  • Mistral Small 3 Input usage
  • Price per 1M tokens

Mistral Small 3 (Output)

Per one-time

$0.8
  • Mistral Small 3 Output usage
  • Price per 1M tokens

Mixtral 8x7B Instruct v0.1 (Input)

Per one-time

$0.6
  • Mixtral 8x7B Instruct v0.1 Input usage
  • Price per 1M tokens

Mixtral 8x7B Instruct v0.1 (Output)

Per one-time

$0.6
  • Mixtral 8x7B Instruct v0.1 Output usage
  • Price per 1M tokens

Marin 8B Instruct (Input)

Per one-time

$0.18
  • Marin 8B Instruct Input usage
  • Price per 1M tokens

Marin 8B Instruct (Output)

Per one-time

$0.18
  • Marin 8B Instruct Output usage
  • Price per 1M tokens

Arcee AI AFM-4.5B (Input)

Per one-time

$0.1
  • Arcee AI AFM-4.5B Input usage
  • Price per 1M tokens

Arcee AI AFM-4.5B (Output)

Per one-time

$0.4
  • Arcee AI AFM-4.5B Output usage
  • Price per 1M tokens

Arcee AI Coder-Large (Input)

Per one-time

$0.5
  • Arcee AI Coder-Large Input usage
  • Price per 1M tokens

Arcee AI Coder-Large (Output)

Per one-time

$0.8
  • Arcee AI Coder-Large Output usage
  • Price per 1M tokens

Arcee AI Maestro (Input)

Per one-time

$0.9
  • Arcee AI Maestro Input usage
  • Price per 1M tokens

Arcee AI Maestro (Output)

Per one-time

$3.3
  • Arcee AI Maestro Output usage
  • Price per 1M tokens

Arcee AI Virtuoso-Large (Input)

Per one-time

$0.75
  • Arcee AI Virtuoso-Large Input usage
  • Price per 1M tokens

Arcee AI Virtuoso-Large (Output)

Per one-time

$1.2
  • Arcee AI Virtuoso-Large Output usage
  • Price per 1M tokens

Cogito v2 preview - 109B MoE (Input)

Per one-time

$0.18
  • Cogito v2 preview - 109B MoE Input usage
  • Price per 1M tokens

Cogito v2 preview - 109B MoE (Output)

Per one-time

$0.59
  • Cogito v2 preview - 109B MoE Output usage
  • Price per 1M tokens

Cogito v2 preview - 405B (Input)

Per one-time

$3.5
  • Cogito v2 preview - 405B Input usage
  • Price per 1M tokens

Cogito v2 preview - 405B (Output)

Per one-time

$3.5
  • Cogito v2 preview - 405B Output usage
  • Price per 1M tokens

Cogito v2 preview - 671B MoE (Input)

Per one-time

$1.25
  • Cogito v2 preview - 671B MoE Input usage
  • Price per 1M tokens

Cogito v2 preview - 671B MoE (Output)

Per one-time

$1.25
  • Cogito v2 preview - 671B MoE Output usage
  • Price per 1M tokens

Cogito v2 preview - 70B (Input)

Per one-time

$0.88
  • Cogito v2 preview - 70B Input usage
  • Price per 1M tokens

Cogito v2 preview - 70B (Output)

Per one-time

$0.88
  • Cogito v2 preview - 70B Output usage
  • Price per 1M tokens

Refuel LLM-2 (Input)

Per one-time

$0.6
  • Refuel LLM-2 Input usage
  • Price per 1M tokens

Refuel LLM-2 (Output)

Per one-time

$0.6
  • Refuel LLM-2 Output usage
  • Price per 1M tokens

Refuel LLM-2 Small (Input)

Per one-time

$0.2
  • Refuel LLM-2 Small Input usage
  • Price per 1M tokens

Refuel LLM-2 Small (Output)

Per one-time

$0.2
  • Refuel LLM-2 Small Output usage
  • Price per 1M tokens

Typhoon 2 70B Instruct (Input)

Per one-time

$0.88
  • Typhoon 2 70B Instruct Input usage
  • Price per 1M tokens

Typhoon 2 70B Instruct (Output)

Per one-time

$0.88
  • Typhoon 2 70B Instruct Output usage
  • Price per 1M tokens

gemma-3n-E4B-it (Input)

Per one-time

$0.02
  • gemma-3n-E4B-it Input usage
  • Price per 1M tokens

gemma-3n-E4B-it (Output)

Per one-time

$0.04
  • gemma-3n-E4B-it Output usage
  • Price per 1M tokens

FLUX.1 Krea [dev]

Per one-time

$0.025
  • FLUX.1 Krea [dev] image generation
  • Price per MP
  • Default steps: 28

FLUX.1 Kontext [dev]

Per one-time

$0.025
  • FLUX.1 Kontext [dev] image generation
  • Price per MP
  • Default steps: 28

FLUX.1 Kontext [pro]

Per one-time

$0.04
  • FLUX.1 Kontext [pro] image generation
  • Price per MP
  • Default steps: 28

FLUX.1 Kontext [max]

Per one-time

$0.08
  • FLUX.1 Kontext [max] image generation
  • Price per MP
  • Default steps: 28

FLUX1.1 [pro]

Per one-time

$0.04
  • FLUX1.1 [pro] image generation
  • Price per MP

FLUX.1 [dev]

Per one-time

$0.025
  • FLUX.1 [dev] image generation
  • Price per MP
  • Default steps: 28

FLUX.1 [pro]

Per one-time

$0.05
  • FLUX.1 [pro] image generation
  • Price per MP
  • Default steps: 28

FLUX.1 [schnell]

Per one-time

$0.0027
  • FLUX.1 [schnell] image generation
  • Price per MP
  • Default steps: 4

FLUX.1 Canny [pro]

Per one-time

$0.05
  • FLUX.1 Canny [pro] image generation
  • Price per MP

Google Imagen 4.0 Preview

Per one-time

$0.04
  • Google Imagen 4.0 Preview image generation
  • Price per MP

Google Imagen 4.0 Fast

Per one-time

$0.02
  • Google Imagen 4.0 Fast image generation
  • Price per MP

Google Imagen 4.0 Ultra

Per one-time

$0.06
  • Google Imagen 4.0 Ultra image generation
  • Price per MP

Gemini Flash Image 2.5 (Nano Banana)

Per one-time

$0.039
  • Gemini Flash Image 2.5 (Nano Banana) image generation
  • Price per MP

ByteDance Seedream 3.0

Per one-time

$0.018
  • ByteDance Seedream 3.0 image generation
  • Price per MP

ByteDance Seedream 4.0

Per one-time

$0.03
  • ByteDance Seedream 4.0 image generation
  • Price per MP

ByteDance SeedEdit

Per one-time

$0.03
  • ByteDance SeedEdit image generation
  • Price per MP

Qwen Image Edit

Per one-time

$0.0032
  • Qwen Image Edit image generation
  • Price per MP

Qwen Image

Per one-time

$0.0058
  • Qwen Image image generation
  • Price per MP

Juggernaut Pro Flux by RunDiffusion

Per one-time

$0.0049
  • Juggernaut Pro Flux by RunDiffusion image generation
  • Price per MP

Juggernaut Lightning Flux by RunDiffusion

Per one-time

$0.0017
  • Juggernaut Lightning Flux by RunDiffusion image generation
  • Price per MP

HiDream-I1-Full

Per one-time

$0.009
  • HiDream-I1-Full image generation
  • Price per MP

HiDream-I1-Dev

Per one-time

$0.0045
  • HiDream-I1-Dev image generation
  • Price per MP

HiDream-I1-Fast

Per one-time

$0.0032
  • HiDream-I1-Fast image generation
  • Price per MP

Ideogram 3.0

Per one-time

$0.06
  • Ideogram 3.0 image generation
  • Price per MP

Dreamshaper

Per one-time

$0.0006
  • Dreamshaper image generation
  • Price per MP

SD XL

Per one-time

$0.0019
  • SD XL image generation
  • Price per MP

Stable Diffusion 3

Per one-time

$0.0019
  • Stable Diffusion 3 image generation
  • Price per MP

Cartesia Sonic-2

Per one-time

$65
  • Cartesia Sonic-2 speech synthesis/processing
  • Price per 1M characters

MiniMax 01 Director (720p/5s)

Per one-time

$0.28
  • MiniMax 01 Director video generation
  • Price per video (720p/5s)

MiniMax Hailuo 02 (768p/10s)

Per one-time

$0.56
  • MiniMax Hailuo 02 video generation
  • Price per video (768p/10s)

MiniMax Hailuo 02 (1080p/6s)

Per one-time

$0.49
  • MiniMax Hailuo 02 video generation
  • Price per video (1080p/6s)

Google Veo 2.0 (720p/5s)

Per one-time

$2.5
  • Google Veo 2.0 video generation
  • Price per video (720p/5s)

Google Veo 3.0 (720p/8s)

Per one-time

$1.6
  • Google Veo 3.0 video generation
  • Price per video (720p/8s)

Google Veo 3.0 + Audio (720p/8s with audio)

Per one-time

$3.2
  • Google Veo 3.0 + Audio video generation
  • Price per video (720p/8s with audio)

Google Veo 3.0 Fast (1080p/8s)

Per one-time

$0.8
  • Google Veo 3.0 Fast video generation
  • Price per video (1080p/8s)

Google Veo 3.0 Fast + Audio (1080p/8s with audio)

Per one-time

$1.2
  • Google Veo 3.0 Fast + Audio video generation
  • Price per video (1080p/8s with audio)

ByteDance Seedance 1.0 Lite (720p/5s)

Per one-time

$0.14
  • ByteDance Seedance 1.0 Lite video generation
  • Price per video (720p/5s)

ByteDance Seedance 1.0 Pro (1080p/5s)

Per one-time

$0.57
  • ByteDance Seedance 1.0 Pro video generation
  • Price per video (1080p/5s)

PixVerse v5 (1080p/5s)

Per one-time

$0.3
  • PixVerse v5 video generation
  • Price per video (1080p/5s)

Kling 2.1 Master (1080p/5s)

Per one-time

$0.92
  • Kling 2.1 Master video generation
  • Price per video (1080p/5s)

Kling 2.1 Standard (720p/5s)

Per one-time

$0.18
  • Kling 2.1 Standard video generation
  • Price per video (720p/5s)

Kling 2.1 Pro (1080p/5s)

Per one-time

$0.32
  • Kling 2.1 Pro video generation
  • Price per video (1080p/5s)

Kling 2.0 Master (1080p/5s)

Per one-time

$0.92
  • Kling 2.0 Master video generation
  • Price per video (1080p/5s)

Kling 1.6 Standard (720p/5s)

Per one-time

$0.19
  • Kling 1.6 Standard video generation
  • Price per video (720p/5s)

Kling 1.6 Pro (1080p/5s)

Per one-time

$0.32
  • Kling 1.6 Pro video generation
  • Price per video (1080p/5s)

Wan 2.2 I2V (720p/5s)

Per one-time

$0.31
  • Wan 2.2 I2V video generation
  • Price per video (720p/5s)

Wan 2.2 T2V (720p/8s)

Per one-time

$0.66
  • Wan 2.2 T2V video generation
  • Price per video (720p/8s)

Vidu 2.0 (720p/8s)

Per one-time

$0.28
  • Vidu 2.0 video generation
  • Price per video (720p/8s)

Vidu Q1 (1080p/5s)

Per one-time

$0.22
  • Vidu Q1 video generation
  • Price per video (1080p/5s)

Sora 2 (720p/8s)

Per one-time

$0.8
  • Sora 2 video generation
  • Price per video (720p/8s)

Sora 2 Pro (720p/8s)

Per one-time

$2.4
  • Sora 2 Pro video generation
  • Price per video (720p/8s)

Sora 2 Pro (1080p/8s)

Per one-time

$4
  • Sora 2 Pro video generation
  • Price per video (1080p/8s)

Whisper Large v3

Per one-time

$0.0015
  • Whisper Large v3 automatic speech recognition
  • Price per audio minute

BGE-Base-EN v1.5

Per one-time

$0.01
  • BGE-Base-EN v1.5 vector embeddings
  • Price per 1M tokens

BGE-Large-EN v1.5

Per one-time

$0.02
  • BGE-Large-EN v1.5 vector embeddings
  • Price per 1M tokens

GTE ModernBERT base

Per one-time

$0.08
  • GTE ModernBERT base vector embeddings
  • Price per 1M tokens

Multilingual e5 large instruct

Per one-time

$0.02
  • Multilingual e5 large instruct vector embeddings
  • Price per 1M tokens

M2-BERT 80M 32K Retrieval

Per one-time

$0.01
  • M2-BERT 80M 32K Retrieval vector embeddings
  • Price per 1M tokens

Mxbai Rerank Large V2

Per one-time

$0.1
  • Mxbai Rerank Large V2 search relevance reranking
  • Price per 1M tokens

Salesforce Llama Rank V1 (8B)

Per one-time

$0.1
  • Salesforce Llama Rank V1 (8B) search relevance reranking
  • Price per 1M tokens

VirtueGuard Text Lite

Per one-time

$0.2
  • VirtueGuard Text Lite content filtering and classification
  • Price per 1M tokens

Llama Guard 4 12B

Per one-time

$0.2
  • Llama Guard 4 12B content filtering and classification
  • Price per 1M tokens

Llama Guard 3 11B Vision Turbo

Per one-time

$0.18
  • Llama Guard 3 11B Vision Turbo content filtering and classification
  • Price per 1M tokens

Llama Guard 3 8B

Per one-time

$0.2
  • Llama Guard 3 8B content filtering and classification
  • Price per 1M tokens

Llama Guard 2 8B

Per one-time

$0.2
  • Llama Guard 2 8B content filtering and classification
  • Price per 1M tokens

Dedicated Endpoint - 1x H200 141GB

Per one-time

$4.99
  • Guaranteed performance
  • Support for custom models
  • Autoscaling & traffic spike handling
  • Hardware: 1x H200 141GB
  • Price per hour

Dedicated Endpoint - 1x H100 80GB

Per one-time

$3.36
  • Guaranteed performance
  • Support for custom models
  • Autoscaling & traffic spike handling
  • Hardware: 1x H100 80GB
  • Price per hour

Dedicated Endpoint - 1x A100 SXM 80GB

Per one-time

$2.56
  • Guaranteed performance
  • Support for custom models
  • Autoscaling & traffic spike handling
  • Hardware: 1x A100 SXM 80GB
  • Price per hour

Dedicated Endpoint - 1x A100 SXM 40GB

Per one-time

$2.4
  • Guaranteed performance
  • Support for custom models
  • Autoscaling & traffic spike handling
  • Hardware: 1x A100 SXM 40GB
  • Price per hour

Dedicated Endpoint - 1x A100 PCIe 80GB

Per one-time

$2.4
  • Guaranteed performance
  • Support for custom models
  • Autoscaling & traffic spike handling
  • Hardware: 1x A100 PCIe 80GB
  • Price per hour

Dedicated Endpoint - 1x L40S 48GB

Per one-time

$2.1
  • Guaranteed performance
  • Support for custom models
  • Autoscaling & traffic spike handling
  • Hardware: 1x L40S 48GB
  • Price per hour

Fine-tuning Up to 16B - Supervised Fine-Tuning LoRA

Per one-time

$0.48
  • Supervised Fine-Tuning LoRA (Up to 16B model size)
  • Price per token processed

Fine-tuning Up to 16B - Direct Preference Optimization LoRA

Per one-time

$0.54
  • Direct Preference Optimization LoRA (Up to 16B model size)
  • Price per token processed

Fine-tuning Up to 16B - Supervised Full Fine-Tuning

Per one-time

$1.2
  • Supervised Full Fine-Tuning (Up to 16B model size)
  • Price per token processed

Fine-tuning Up to 16B - Direct Preference Optimization Full Fine-Tuning

Per one-time

$1.35
  • Direct Preference Optimization Full Fine-Tuning (Up to 16B model size)
  • Price per token processed

Fine-tuning 17B-69B - Supervised Fine-Tuning LoRA

Per one-time

$1.5
  • Supervised Fine-Tuning LoRA (17B-69B model size)
  • Price per token processed

Fine-tuning 17B-69B - Direct Preference Optimization LoRA

Per one-time

$1.65
  • Direct Preference Optimization LoRA (17B-69B model size)
  • Price per token processed

Fine-tuning 17B-69B - Supervised Full Fine-Tuning

Per one-time

$3.75
  • Supervised Full Fine-Tuning (17B-69B model size)
  • Price per token processed

Fine-tuning 17B-69B - Direct Preference Optimization Full Fine-Tuning

Per one-time

$4.12
  • Direct Preference Optimization Full Fine-Tuning (17B-69B model size)
  • Price per token processed

Fine-tuning 70-100B - Supervised Fine-Tuning LoRA

Per one-time

$2.9
  • Supervised Fine-Tuning LoRA (70-100B model size)
  • Price per token processed

Fine-tuning 70-100B - Direct Preference Optimization LoRA

Per one-time

$3.2
  • Direct Preference Optimization LoRA (70-100B model size)
  • Price per token processed

Fine-tuning 70-100B - Supervised Full Fine-Tuning

Per one-time

$7.25
  • Supervised Full Fine-Tuning (70-100B model size)
  • Price per token processed

Fine-tuning 70-100B - Direct Preference Optimization Full Fine-Tuning

Per one-time

$8
  • Direct Preference Optimization Full Fine-Tuning (70-100B model size)
  • Price per token processed

Specialized Fine-tuning gpt-oss-120B - Supervised Fine-Tuning LoRA

Per one-time

$5
  • Supervised Fine-Tuning LoRA for gpt-oss-120B model
  • Limited to LoRA fine-tuning
  • Minimum charge: $6.00
  • Price per token processed

Specialized Fine-tuning gpt-oss-120B - Direct Preference Optimization LoRA

Per one-time

$12.5
  • Direct Preference Optimization LoRA for gpt-oss-120B model
  • Limited to LoRA fine-tuning
  • Minimum charge: $6.00
  • Price per token processed

Specialized Fine-tuning Llama 4 Scout Instruct - Supervised Fine-Tuning LoRA

Per one-time

$3
  • Supervised Fine-Tuning LoRA for Llama 4 Scout Instruct model
  • Limited to LoRA fine-tuning
  • Minimum charge: $6.00
  • Price per token processed

Specialized Fine-tuning Llama 4 Scout Instruct - Direct Preference Optimization LoRA

Per one-time

$7.5
  • Direct Preference Optimization LoRA for Llama 4 Scout Instruct model
  • Limited to LoRA fine-tuning
  • Minimum charge: $6.00
  • Price per token processed

Specialized Fine-tuning Llama 4 Maverick Instruct - Supervised Fine-Tuning LoRA

Per one-time

$8
  • Supervised Fine-Tuning LoRA for Llama 4 Maverick Instruct model
  • Limited to LoRA fine-tuning
  • Minimum charge: $16.00
  • Price per token processed

Specialized Fine-tuning Llama 4 Maverick Instruct - Direct Preference Optimization LoRA

Per one-time

$20
  • Direct Preference Optimization LoRA for Llama 4 Maverick Instruct model
  • Limited to LoRA fine-tuning
  • Minimum charge: $16.00
  • Price per token processed

Specialized Fine-tuning DeepSeek Models - Supervised Fine-Tuning LoRA

Per one-time

$10
  • Supervised Fine-Tuning LoRA for DeepSeek-R1, DeepSeek-R1-0528, DeepSeek-V3, DeepSeek-V3-0324, DeepSeek-V3.1, DeepSeek-V3.1-Base models
  • Limited to LoRA fine-tuning
  • Minimum charge: $20.00
  • Price per token processed

Specialized Fine-tuning DeepSeek Models - Direct Preference Optimization LoRA

Per one-time

$25
  • Direct Preference Optimization LoRA for DeepSeek-R1, DeepSeek-R1-0528, DeepSeek-V3, DeepSeek-V3-0324, DeepSeek-V3.1, DeepSeek-V3.1-Base models
  • Limited to LoRA fine-tuning
  • Minimum charge: $20.00
  • Price per token processed

Specialized Fine-tuning Qwen3-Coder-480B-A35B-Instruct - Supervised Fine-Tuning LoRA

Per one-time

$9
  • Supervised Fine-Tuning LoRA for Qwen3-Coder-480B-A35B-Instruct model
  • Limited to LoRA fine-tuning
  • Minimum charge: $18.00
  • Price per token processed

Specialized Fine-tuning Qwen3-Coder-480B-A35B-Instruct - Direct Preference Optimization LoRA

Per one-time

$22.5
  • Direct Preference Optimization LoRA for Qwen3-Coder-480B-A35B-Instruct model
  • Limited to LoRA fine-tuning
  • Minimum charge: $18.00
  • Price per token processed

Specialized Fine-tuning Qwen3-235B-A22B Models - Supervised Fine-Tuning LoRA

Per one-time

$6
  • Supervised Fine-Tuning LoRA for Qwen3-235B-A22B, Qwen3-235B-A22B-Instruct-2507 models
  • Limited to LoRA fine-tuning
  • No minimum charge
  • Price per token processed

Specialized Fine-tuning Qwen3-235B-A22B Models - Direct Preference Optimization LoRA

Per one-time

$15
  • Direct Preference Optimization LoRA for Qwen3-235B-A22B, Qwen3-235B-A22B-Instruct-2507 models
  • Limited to LoRA fine-tuning
  • No minimum charge
  • Price per token processed

Code Sandbox (vCPU)

Per one-time

$0.0446
  • Code Sandbox vCPU usage
  • Price per vCPU per hour
  • Customize VM sandboxes

Code Sandbox (GiB RAM)

Per one-time

$0.0149
  • Code Sandbox GiB RAM usage
  • Price per GiB RAM per hour
  • Customize VM sandboxes

Code Interpreter Session

Per one-time

$0.03
  • Code Interpreter session (60 minutes)
  • Price per session

Instant Cluster - NVIDIA HGX H100 SXM

Per one-time

$2.99
  • Ready to use, self-service GPUs
  • NVIDIA HGX H100 SXM
  • Price per hour per GPU

Instant Cluster - NVIDIA HGX H200

Per one-time

$3.79
  • Ready to use, self-service GPUs
  • NVIDIA HGX H200
  • Price per hour per GPU

Instant Cluster - NVIDIA HGX B200

Per one-time

$5.5
  • Ready to use, self-service GPUs
  • NVIDIA HGX B200
  • Price per hour per GPU

Reserved Cluster - NVIDIA GB200 NVL72 384GB HBM3e

Per one-time

Free
  • Dedicated capacity
  • Expert support
  • Hardware: NVIDIA GB200 NVL72
  • GPU Memory: 384GB HBM3e
  • Price per hour

Reserved Cluster - NVIDIA B200 192GB HBM3e

Per one-time

Free
  • Dedicated capacity
  • Expert support
  • Hardware: NVIDIA B200
  • GPU Memory: 192GB HBM3e
  • Price per hour

Reserved Cluster - NVIDIA H200 141GB HBM3e

Per one-time

$2.09
  • Dedicated capacity
  • Expert support
  • Hardware: NVIDIA H200
  • GPU Memory: 141GB HBM3e
  • Starting at price per hour

Reserved Cluster - NVIDIA H100 80GB HBM2e

Per one-time

$1.75
  • Dedicated capacity
  • Expert support
  • Hardware: NVIDIA H100
  • GPU Memory: 80GB HBM2e
  • Starting at price per hour

Reserved Cluster - NVIDIA A100 80GB HBM2e

Per one-time

$1.3
  • Dedicated capacity
  • Expert support
  • Hardware: NVIDIA A100
  • GPU Memory: 80GB HBM2e
  • Starting at price per hour

Frontier AI Factory

Per one-time

Free
  • Large-scale, custom-built private GPU clusters
  • NVIDIA Blackwell GPUs at scale
  • Custom quote required

Shared Filesystem Storage

Per monthly

$0.16
  • High-bandwidth, parallel filesystem
  • Colocated with compute
  • Price per GiB per month

✓ Free plan • ✓ Plans from $0.0006 / one-time

Integrations

OpenAI
DeepSeek
Qwen
Llama
Kimi K2
Apriel

Want your AI tool listed here?

Start with a free eligibility check.

Submit