Forgeworker guide

AI APIs and model infrastructure

Compare AI APIs for language, image, audio, video, embeddings, and custom models by capabilities, access, pricing source, and deployment model.

An AI API should be chosen for the production job around the model: reliability, latency, output controls, observability, data policy, price, and a clean failure path. This directory links to official providers and separates verified canaries from untested claims.

Direct provider or model router

A direct provider can offer the tightest access to its models and features. A router can simplify switching and comparison across providers. Infrastructure platforms can host public, custom, or containerized models when you need more control.

Price the complete request

For language models, count input, cached input, output, tools, and retries. For media, include resolution, duration, enhancement, storage, and failed generations. Read the linked live pricing source because provider rates and free tiers can change.

Design for failure before launch

Set timeouts, idempotency keys, spend ceilings, rate limits, logging, cancellation, and a useful fallback. Do not expose a provider key in browser code. Test success, refusal, timeout, provider error, and billing settlement as separate states.

Questions to ask before choosing

Does Forgeworker resell every API in this directory?

No. External API listings link to the official provider. ForgeMeter access is only claimed when a Forgeworker integration and production canary are explicitly shown.

How should I compare AI API prices?

Use one representative workload and count every billable input, output, media, tool, retry, and storage component. Then compare observed quality and latency alongside cost.

Cloudflare Workers AI

Run serverless AI models close to users. Cloudflare’s hosted model platform for text, image, speech, embeddings, and other inference workloads.

OpenAI API

Multimodal models and agent-building APIs. APIs for text, reasoning, images, audio, embeddings, moderation, and agentic application workflows.

Anthropic API

Claude models for reasoning, coding, and agents. A hosted API for Claude models with text, vision, tool use, prompt caching, and agent workflows.

Google Gemini API

Google’s multimodal model API and developer studio. Gemini models for text, images, audio, video understanding, embeddings, and tool-enabled applications.

OpenRouter

One API for a wide catalog of language models. A model routing API that provides a common interface across many hosted language-model providers.

Together AI

Hosted open-model inference and fine-tuning. A cloud platform for running, fine-tuning, and deploying open generative models.

Replicate

Run public and custom models through an API. A hosted model platform for running community models and deploying custom inference versions.

Runware

Fast image and media inference through one API. A hosted inference service focused on image generation, editing, custom models, and media workflows.

RunPod Serverless

Scale custom GPU workers down to zero. GPU serverless infrastructure for custom containers, private models, ComfyUI workflows, and specialized inference.

GroqCloud

Low-latency inference for supported language models. A hosted inference API using Groq hardware for supported text, speech, and multimodal models.

Mistral Studio

Mistral models, agents, and document tools. Mistral’s developer workspace for model APIs, agents, usage tracking, document understanding, and deployment…

Perplexity API

Search-grounded answers through an API. APIs for web-grounded search, answer generation, and research-oriented application experiences.

This server-rendered summary is available to search engines and no-JavaScript visitors. JavaScript loads the full interactive experience.