GPT Transcribe is OpenAI's speech-to-text model for completed audio files, streamed file transcripts, and committed turns in Realtime sessions over WebSocket. It supports unstructured context, keyword hints, and multiple language hints to improve transcription of domain terms, multilingual audio, and code-switching, and reports which languages it actually detected. On OpenAI's Common Voice benchmark across 22 languages, it cuts word error rate to 19.27% against 40.37% for whisper-1. It returns plain text only, without subtitle formats, word/segment timestamps, or speaker labels.
Get started by making your first request to FastRouter. Use your favorite SDK or make direct API calls.
<FASTROUTER_API_KEY> with your actual API keyGPT-6 Astra is OpenAI’s flagship frontier model for complex end-to-end work, featuring very long context (up to ~1.05M tokens), advanced reasoning with configurable effort levels, and strong performance on coding, computer use, research, and professional tasks. It supports text and image inputs with text-only outputs, tool use via the Responses API (including web search, code execution, file search, and computer use), and structured outputs. Astra is heavily safety-aligned and designed for multi-step workflows, async tool calling, and mid-turn steering in demanding enterprise and developer applications.
GPT-5.3-Codex is OpenAI's most capable agentic coding model to date, combining frontier coding performance with general work automation capabilities—enabling developers and professionals to automate complex tasks across the entire software lifecycle, from code writing to infrastructure management and cybersecurity.
GPT Image 2.5 Flare is OpenAI's fast text-to-image generation and editing model, delivering higher-quality images than GPT Image 2 at up to 50% lower latency. It is the default choice for most API applications, sharing Images 2.5's improvements in reference fidelity, precise edits, and multi-turn consistency, and is well suited to creator and social content, product experiences, visual search, rapid prototyping, and high-volume generation.
gpt-oss-120b is OpenAI's larger open-weight model, designed for powerful reasoning, agentic tasks, and versatile developer use cases. It offers native function calling, python tool calls, structured outputs, full chain-of-thought access, and configurable reasoning effort (low/medium/high), released under the Apache 2.0 license. The MoE weights are post-trained with MXFP4 quantization for efficient deployment.
GPT-5.6 Luna is OpenAI’s fast, cost-efficient tier in the GPT‑5.6 family, optimized for low-latency, high-volume workloads such as drafting, summarization, and routine automation. It supports large-context multimodal chat and tool use with a reported 1.5M-token context window and standard OpenAI function-calling/tooling semantics. Luna trades some peak capability for significantly lower cost compared to Sol and Terra, making it suitable as a default production workhorse model.
GPT Image 2.5 Sunburst is OpenAI's high-fidelity text-to-image generation and editing model, offering an extra level of precision for detailed creative work at the cost of longer generation times than Flare. It shares Images 2.5's improvements in reference fidelity, precise edits, and multi-turn consistency, with sharper detail, more natural lighting, richer textures, and better handling of complex layouts and transparent backgrounds. Recommended for premium visual workflows like production-ready campaign creative or polished product imagery.
OpenAI's gpt-image-2 (also referred to as GPT Image 2 or ChatGPT Images 2.0) is a state-of-the-art text-to-image generation and editing model that transforms natural-language prompts into high-quality, photorealistic visuals with exceptional prompt fidelity, accurate text rendering, and advanced reasoning capabilities.
OpenAI’s gpt-image-1.5 is a next‑generation image generation and editing model designed to create production‑quality visuals from text prompts and to edit existing images with fine control.
GPT-3.5 Turbo represents OpenAI's efficiency-optimized language model designed for maximum responsiveness in conversational and completion applications. This solution delivers optimal performance for real-time interactions, code generation, and content creation tasks with minimal latency. Despite its focus on processing speed, the model maintains impressive capabilities in understanding and generating both natural language and programming code across diverse domains. GPT-3.5 Turbo is specifically tuned for dialogue applications, making it particularly effective for chatbots, virtual assistants, and interactive systems requiring rapid responses. With training data extending to September 2021, the model possesses comprehensive general knowledge while maintaining computational efficiency suitable for high-throughput applications. As OpenAI's most cost-effective and responsive production model, it provides an optimal balance of capability, performance, and affordability for applications where processing speed and deployment economics are critical considerations.
GPT-3.5 Turbo has been largely superseded by more advanced models in OpenAI's lineup. While still available for legacy applications and cost-sensitive implementations, its capabilities are now significantly outpaced by GPT-4o and the O-series models. The model remains suitable for basic conversational AI, content generation, and simple function-calling scenarios, but lacks the advanced reasoning, multilingual capabilities, and multimodal features of newer models. For applications requiring only core language processing without the need for cutting-edge performance, it continues to offer a cost-effective option.
An older GPT-3.5 Turbo model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Training data: up to Sep 2021.
This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up to Sep 2021.
The original GPT-4 model has been replaced by more advanced versions in OpenAI's lineup. With knowledge cutoff from September 2021, it lacks awareness of more recent events and developments. GPT-4o now serves as OpenAI's primary multimodal model, offering similar capabilities with substantially improved speed, cost efficiency, and multilingual support. For applications requiring more advanced reasoning, the O-series models (o1, o3, o4) provide specialized capabilities for complex problem-solving in mathematics, science, coding, and other technical domains.
GPT-4 Turbo has been superseded by GPT-4o as OpenAI's flagship multimodal model. While it still offers vision capabilities, JSON mode, and function calling, GPT-4o provides faster response times (2x faster), enhanced multilingual support, and improved visual understanding at 50% reduced cost. As of May 2025, GPT-4 Turbo is being phased out in favor of GPT-4o and the newer reasoning-focused O-series models.
GPT-4.1 is OpenAI's latest flagship model showcasing significant advances in coding, instruction following, and long-context understanding. It processes up to 1 million tokens (approximately 750,000 words) in a single context window, with knowledge updated to June 2024. Performance benchmarks demonstrate substantial improvements over GPT-4o, including 54.6% completion on SWE-bench Verified coding tasks (21.4% increase) and 38.3% on MultiChallenge instruction following (10.5% increase). The model excels at real-world software engineering with reduced extraneous edits (from 9% to 2%), superior code exploration capabilities, and enhanced long-document comprehension. GPT-4.1 provides improved agentic reliability while maintaining a lower price point than its predecessors, making it particularly valuable for development environments, document analysis, and enterprise knowledge systems.
GPT-4.1 Mini delivers GPT-4o-level performance with significantly improved efficiency metrics across latency and cost. This mid-sized model maintains the full 1 million token context window of its larger counterpart while achieving impressive benchmark scores: 45.1% on hard instruction evaluations, 35.8% on MultiChallenge, and 84.1% on IFEval. Despite its reduced parameter count, GPT-4.1 Mini demonstrates robust coding capabilities (31.6% on Aider's polyglot diff benchmark) and strong vision understanding. The model reduces latency by nearly half and costs 83% less than GPT-4o, making it ideal for interactive applications with tight performance constraints, high-throughput services, and cost-sensitive enterprise deployments requiring responsive AI capabilities without compromising on quality.
GPT-4.1 Nano is OpenAI's most efficient model optimized for maximum speed and minimum cost in the GPT-4.1 family. Released in April 2025, this compact model maintains the full 1 million token context window while delivering impressive benchmark performance: 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding—outperforming even GPT-4o mini. At just $0.10 per million input tokens and $0.40 per million output tokens, it represents OpenAI's most affordable option. The model excels in latency-critical applications including classification, autocompletion, information extraction, and high-throughput document processing where efficiency and cost considerations are paramount.
GPT-4o (o for omni) is OpenAI's multimodal flagship model combining advanced language processing with sophisticated visual understanding. Released in 2024, this versatile solution processes both text and image inputs to generate high-quality text outputs across diverse applications. While maintaining the same intelligence level as GPT-4 Turbo, it delivers twice the processing speed and 50% greater cost efficiency. The model features significantly enhanced multilingual capabilities and improved visual analysis functions, making it particularly valuable for global applications requiring both text and image processing. As the foundation for OpenAI's consumer and enterprise offerings, GPT-4o provides a balanced combination of quality, speed, and affordability that makes advanced AI capabilities more accessible for diverse use cases.
GPT-4o has established itself as OpenAI's primary multimodal model, integrating text, images, and audio processing in a single unified architecture. Since its initial release, it has received updates that further improve its rapid response capabilities (averaging 320ms), enhance its non-English language performance, and strengthen its visual understanding. Its multimodal design allows for natural, intuitive interactions across input types, making it particularly effective for applications requiring seamless switching between text, image, and audio comprehension. GPT-4o maintains competitive performance while offering significant speed and cost advantages over specialized models.
GPT-4o (August 2024 version) enhances OpenAI's flagship multimodal model with sophisticated structured output capabilities. This specific release introduces advanced JSON schema support through the response_format parameter, enabling precise control over output structure and data formatting for applications requiring consistent, well-defined response patterns. The model maintains all core GPT-4o capabilities: processing both text and image inputs while generating high-quality text outputs with double the speed and 50% greater cost efficiency than GPT-4 Turbo. Additional refinements include improved non-English language processing and enhanced visual analysis capabilities. Released in August 2024, this version particularly benefits developers building data-driven applications, APIs, and systems requiring predictable, structured information exchange with minimal post-processing. It represents an important evolution in OpenAI's efforts to make their models more useful for programmatic and enterprise applications requiring strict output formats alongside traditional natural language generation.
GPT-4o (November 2024 version) enhances OpenAI's flagship multimodal model with specialized improvements in creative writing and document analysis. This updated release delivers more natural, engaging, and contextually tailored writing with improved relevance and readability across various content types. The model demonstrates significantly enhanced document processing capabilities, providing deeper insights and more comprehensive responses when working with uploaded files. While maintaining the core multimodal architecture that processes both text and image inputs, this version incorporates refinements that improve multilingual processing and visual understanding. The model continues to deliver the performance level of GPT-4 Turbo with double the processing speed and 50% greater cost efficiency, making it particularly valuable for content creation, document analysis, and multilingual applications requiring both quality and responsiveness.
GPT-4o mini (o for omni) is OpenAI's efficient, cost-effective compact model designed for targeted applications. This solution processes both textual and visual inputs while generating text responses (including Structured Outputs). It is particularly well-suited for fine-tuning applications, and outputs from more advanced models like GPT-4o can be condensed to GPT-4o-mini to achieve comparable results with reduced expenses and response times.
GPT-4o mini (July 2024 version) represents OpenAI's efficiency-focused multimodal model combining strong performance with minimal resource requirements. Released in July 2024, this model processes both text and image inputs while generating text outputs at a significantly reduced cost—60% cheaper than GPT-3.5 Turbo at just $0.15 per million input tokens and $0.60 per million output tokens. Despite its optimized size, the model achieves impressive benchmark results including an 82% score on MMLU and 87% on MGSM for mathematical reasoning, outperforming comparable small models like Gemini 1.5 Flash and Claude 3 Haiku. With a 128K token context window and exceptionally fast processing speed (202 tokens per second), this model is particularly valuable for high-throughput applications, real-time services, and cost-sensitive deployments requiring responsive multimodal capabilities without sacrificing core intelligence.
GPT‑5 is OpenAI’s flagship model, tailored for complex, multi-step reasoning and high-fidelity code generation. It combines superior world knowledge with streamlined “agentic” capabilities, enabling it to autonomously tackle tasks with minimal prompting. The model excels in logic-driven scenarios, debugging, and structured workflows—making it ideal for developers and professionals with high demands. With enhanced reasoning and smoother interaction, GPT‑5 offers users both intelligence and autonomy in one package.
GPT‑5 Mini offers a lighter and more cost-effective option, designed for everyday chat and instruction-following use cases. It retains much of GPT‑5’s reasoning power while running efficiently at lower compute costs, striking a balance between performance and affordability. Perfect for startups, pupils, or anyone mindful of resource budget, GPT‑5 Mini makes advanced capabilities more accessible. Despite its compact size, it still packs solid intelligence for routine workflows and interactions.
GPT‑5 Nano is optimized for ultra-fast performance and very low latency—ideal for real-time or embedded applications. Stripped down for speed, it excels at simple instruction-following and classification tasks, making it perfect for high-throughput API usage. With minimal overhead and rapid response times, GPT‑5 Nano is a great fit for lightweight environments or devices with limited compute capabilities. Yet, it preserves sufficient reasoning capability to remain practical across many practical scenarios.
GPT-5.1, released in November 2025, is the latest flagship large language model from OpenAI, featuring significant upgrades in intelligence, reasoning, and user experience compared to its predecessor, GPT-5.
GPT-5.2 is OpenAI’s flagship frontier model in the GPT‑5 series, designed for the most demanding reasoning, coding, and long‑context workloads. It emphasizes stronger general intelligence, more reliable instruction following, and advanced safety compared with earlier 5.1 releases.
GPT-5.2 Pro is OpenAI's flagship enterprise-grade reasoning model released on December 10, 2025, designed for sophisticated professional workflows requiring advanced problem-solving capabilities.
GPT-5.4 is the newest flagship model from OpenAI, merging the Codex and GPT product lines into one unified system. It supports a context window exceeding 1 million tokens — with up to 922K tokens for input and 128K for output — and accepts both text and image inputs. This allows it to handle high-context reasoning, code generation, and multimodal analysis all within a single workflow. The model shows notable improvements in areas such as coding, document comprehension, tool integration, and instruction adherence. It is built to serve as a reliable default for both general-purpose and software engineering tasks, capable of producing production-ready code, aggregating insights from diverse sources, and carrying out intricate multi-step processes with reduced iterations and improved token efficiency.
GPT-5.4 mini is OpenAI's compact, high-performance model in the GPT-5.4 family, balancing advanced reasoning, multimodal capabilities, and efficiency for high-volume workloads like coding, agent workflows, and production-scale applications. It delivers significant improvements over GPT-5 mini in coding, reasoning, multimodal understanding, and tool use, running over 2x faster while approaching GPT-5.4 performance on benchmarks like SWE-Bench Pro and OSWorld-Verified. The model supports text and image inputs (no audio/video) with text outputs, a 400,000-token context window, up to 128,000 max output tokens, and an August 31, 2025 knowledge cutoff.
GPT-5.4 nano is OpenAI's most lightweight and cost-efficient model in the GPT-5.4 family, optimized for speed-critical, high-volume tasks like classification, data extraction, ranking, and sub-agent execution. It prioritizes low latency and efficiency over deep reasoning, making it ideal for real-time systems, background tasks, distributed agent architectures, coding assistants, and multimodal applications involving images. The model supports text and image inputs (no audio or video) with text outputs, a 400,000-token context window, and up to 128,000 max output tokens; its knowledge cutoff is August 31, 2025.
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs. Optimized for step-by-step reasoning, instruction following, and accuracy, GPT-5.4 Pro excels at agentic coding, long-context workflows, and multi-step problem solving.
OpenAI's GPT-5.5 is a large language model (LLM) released on April 23, 2026, codenamed "Spud," positioned as the company's smartest model yet for complex real-world tasks like coding, research, data analysis, and agentic workflows. It features variants including GPT-5.5 Thinking and GPT-5.5 Pro (not available to free-tier users), with API access starting April 24 after safeguards implementation
OpenAI's GPT-5.5 Pro is a premium variant of the GPT-5.5 model, designed for higher-tier users (Pro, Business, Enterprise) handling demanding workloads like advanced research, business analysis, legal tasks, and deep reasoning with superior accuracy and depth. It builds on GPT-5.5's agentic capabilities but adds extended reasoning modes for complex, multi-step workflows.
GPT-5.6 Sol is OpenAI’s flagship frontier model in the GPT‑5.6 family, offering advanced reasoning, coding, and cybersecurity capabilities with multimodal text and image input and text output. It is released in a limited preview for trusted partners via the OpenAI API, with enhanced safety systems for higher‑risk and sensitive use cases. The model is designed for complex professional work, long‑horizon tasks, and agentic workflows, with improved robustness and prompt caching over prior generations.
GPT-5.6 Terra is the balanced, mid‑tier model in OpenAI’s GPT‑5.6 family, designed for everyday professional and enterprise workloads. It supports text and image inputs with text outputs, offers strong reasoning and coding capabilities comparable to GPT‑5.5 at roughly half the cost, and is available in a limited preview via the OpenAI API. Terra is intended as a cost‑efficient default for broad knowledge work, analysis, and software development tasks.
GPT-image-1 is OpenAI's latest multimodal image generation model, released in April 2025, and represents a significant advancement over previous models like DALL-E. It can generate high-resolution images—up to 4096×4096 pixels—from natural language prompts, with improved fidelity and consistency in handling complex scenes and detailed instructions. Key features include robust text-to-image generation, reliable text rendering within images, and advanced editing capabilities such as image-to-image transformation and inpainting, allowing users to upload and modify existing images using text prompts.
OpenAI's gpt-image-1-mini is a natively multimodal image model optimized for affordability and efficiency, making it suitable for projects that need high-throughput AI image generation at a lower cost. This model is designed to handle both text and image inputs and can generate new images, perform targeted image editing (like inpainting), and use reference images for style or content guidance.
GPT-OSS-120B is OpenAI’s most powerful open-weight model, designed for production-grade reasoning tasks with 117 billion parameters and a Mixture-of-Experts (MoE) architecture that activates 5.1 billion parameters per token. It uses a 36-layer transformer and supports long contexts (up to 128K tokens), excelling in areas like coding, math competitions, and health applications. The model offers strong tool-use capabilities (e.g., browsing, function calling) and performs on par with o4-mini in core reasoning benchmarks.
GPT-OSS-20B is a lightweight, efficient open-weight model optimized for low-latency, on-device, and edge deployments. With 21 billion parameters and a similar MoE structure as its larger sibling, it supports 128K token contexts while running comfortably on 16GB consumer hardware. It matches or beats o3-mini on standard tasks and excels in coding, math, and health-related reasoning.
gpt-oss-20b is OpenAI's smaller open-weight model, designed for lower latency, local, or specialized use cases while retaining strong reasoning and agentic capabilities. It supports native function calling, python tool calls, structured outputs, full chain-of-thought access, and configurable reasoning effort, released under the Apache 2.0 license.
This is OpenAI's general-availability realtime model, capable of responding to audio and text inputs in realtime over WebRTC, WebSocket, or SIP connections. Supports a 32,000 token context window and up to 4,096 output tokens per response (knowledge cutoff: Oct 01, 2023). Also supports image input (text/audio/image → text/audio).
GPT Realtime 1.5 is OpenAI's flagship low-latency, speech-to-speech model optimized for live conversational systems, voice agents, and customer support, using persistent streaming sessions via the Realtime API (WebRTC/WebSocket).
GPT-Realtime-2.1 is OpenAIs realtime reasoning model for complex voice-agent workflows.
A cost-efficient version of GPT Realtime, capable of responding to audio and text inputs in realtime over WebRTC, WebSocket, or SIP connections. Supports a 32,000 token context window and up to 4,096 output tokens per response (knowledge cutoff: Oct 01, 2023). Also supports image input (text/audio/image → text/audio).
OpenAI's o1 model family is its most advanced yet, built to reason more deeply through extended thought and large-scale reinforcement learning. Optimized for STEM tasks, o1 consistently achieves PhD-level accuracy on benchmarks in physics, chemistry, and biology.
OpenAI o3 is an advanced reasoning model, setting new standards in coding, math, science, and visual understanding. It outperforms on benchmarks like Codeforces, SWE-bench (without custom scaffolding), and MMMU. Ideal for complex, nuanced problems, o3 excels at visual tasks and reduces major errors by 20% compared to o1 in real-world evaluations. It’s especially strong in programming, business, and creative ideation, with early users praising its ability to generate and critically assess novel ideas in biology, math, and engineering.
OpenAI o3-mini (January 2025 version) is a specialized reasoning model optimized for STEM domains while maintaining cost efficiency. This model features adjustable reasoning capabilities and demonstrates exceptional performance in science, mathematics, and coding tasks, with expert evaluators preferring its responses 56% of the time over previous versions and noting a 39% reduction in major errors on complex questions. Even at medium reasoning settings, o3-mini matches the larger o1 model's performance on challenging evaluations like AIME and GPQA while maintaining significantly lower latency and cost profiles. The model supports developer-focused features including function calling, structured outputs, and streaming capabilities, though it lacks vision processing functionality. This text-only design focuses computational resources on reasoning quality, making it particularly valuable for technical applications requiring accurate problem-solving without multimodal requirements.
OpenAI o3-mini-high is the enhanced reasoning variant of o3-mini with reasoning_effort permanently set = high for maximum problem-solving capabilities. Released in January 2025, this specialized configuration prioritizes thorough analysis and step-by-step reasoning over processing speed, making it ideal for complex STEM problems requiring in-depth analysis. While maintaining the same core architecture as standard o3-mini, this variant allocates additional computational resources to reasoning processes, resulting in higher accuracy on challenging problems in science, mathematics, and coding domains. The model maintains all developer-focused features of the standard version, including function calling, structured outputs, and streaming capabilities, while focusing its additional reasoning capacity on problem complexity rather than multimodal processing. This configuration is especially valuable for applications where solution quality and reasoning thoroughness outweigh processing speed considerations, such as educational tools, research assistants, and technical problem-solving systems requiring maximum accuracy in areas demanding rigorous analytical thinking.
OpenAI o4-mini combines exceptional efficiency with sophisticated reasoning capabilities in a compact model from the o-series. Released in April 2025, this streamlined solution maintains powerful multimodal and agentic capabilities while significantly reducing computational requirements. The model achieves remarkable benchmark performance, including near-perfect scores on AIME with Python (99.5%) and competitive results on SWE-bench, outperforming its predecessor o3-mini and approaching the larger o3 model in several domains. Despite its optimized size, o4-mini excels in STEM tasks, visual problem-solving (MathVista, MMMU), and code editing applications. Its refined reinforcement learning architecture enables sophisticated capabilities like tool chaining, structured output generation, and multi-step task completion with minimal latency—often resolving complex problems in under a minute, making it ideal for high-throughput scenarios prioritizing both speed and quality.
OpenAI o4-mini-high is the enhanced reasoning variant of o4-mini with reasoning_effort permanently set to "high" for maximum problem-solving capabilities. Released in April 2025, this configuration prioritizes thorough analysis and step-by-step reasoning over processing speed, making it ideal for complex STEM problems, mathematical proofs, and detailed code generation tasks. While maintaining the same model architecture as standard o4-mini, this variant allocates additional computational resources to reasoning processes, resulting in higher accuracy on complex problems with a corresponding increase in token usage and processing time. The model demonstrates exceptional performance on sophisticated benchmarks including AIME and achieves impressive results on visual reasoning tasks. This specialized configuration is particularly valuable for applications where solution quality and reasoning thoroughness outweigh speed considerations, such as educational tools, research assistance, and complex problem-solving systems requiring maximum accuracy.
Omni Moderation (Latest) is a state-of-the-art moderation model designed for real-time classification of unsafe, offensive, or policy-violating content. It supports text moderation across multiple languages and content types, including user messages, document uploads, code, and social media content. The model is optimized for low false positives and fast ruling in both API and streaming workflows.
Sora-2 is OpenAI's state-of-the-art AI video and audio generation model, released in September 2025, and marks a major leap in controllable, realistic, and physically accurate video synthesis. It transforms text prompts and images into cinematic HD videos with synchronized audio, realistic motion, and advanced world simulation—enabling users to generate scenes with perfect lip-sync, sound effects, and genuine continuity.
Sora-2 Pro is OpenAI’s premium offering for its latest AI-powered video and audio generation system, Sora 2, launched in October 2025. Sora-2 Pro is designed for professional creators and advanced users, giving them enhanced creative controls, access to longer and higher-quality generative video clips, and advanced editing tools.
text-embedding-3-large is OpenAI's high-capacity embedding model for text, aimed at delivering the strongest semantic encoding. It produces 3072-dimensional embeddings and achieves the highest accuracy among OpenAI's embedding models (about 64.6% on the MTEB benchmark). Because of its size, it is slower and more costly – around $0.13 per 1K tokens – compared to smaller models. This makes it well-suited to use cases where embedding quality is critical, such as complex document search or analytics. In summary, text-embedding-3-large outperforms both text-embedding-3-small and the older ada-002 model in accuracy at the expense of higher computational cost and latency.
text-embedding-3-small is a compact OpenAI text embedding model designed for general-purpose semantic tasks (search, classification, similarity) with high throughput. It produces 1536-dimensional vectors and is optimized for efficiency. Its cost is very low (about $0.02 per 1K tokens), making it far cheaper than larger models, though its accuracy is correspondingly slightly lower (MTEB benchmark ~62.3%). The model is ideal for cost-sensitive or high-throughput scenarios where good semantic embeddings are needed but ultra-high accuracy is not critical. In practice, it still performs well on search and retrieval tasks, but trades off some accuracy and richness of representation relative to the larger embedding models.
text-embedding-ada-002 (often called ada v2) is an earlier-generation OpenAI embedding model for general semantic tasks like text and code search, clustering, and classification. It produces 1536-dimensional embeddings and was initially state-of-the-art for tasks such as text search, code search, and sentence similarity. Its performance on benchmarks is moderate (around 61% on MTEB), and it costs about $0.10 per 1K tokens. Ada-002 is suited to general retrieval and similarity tasks but is generally considered a legacy model; newer v3 models (3-small and 3-large) have mostly superseded it for better efficiency or accuracy. In fact, documentation notes that ada-002 is primarily kept for legacy use cases, while most new applications use the 3-small or 3-large models for improved performance.
OpenAI: TTS-1 is a text-to-speech model optimised for real-time use cases, converting text into natural-sounding spoken audio with low latency. It ships six built-in voices (alloy, echo, fable, onyx, nova, shimmer), supports six output containers (mp3, opus, aac, flac, wav, pcm) and adjustable speech speed from 0.25x to 4.0x. Input is capped at 4096 characters per request. For higher audio fidelity at double the cost, use tts-1-hd; for voice-style control via natural-language instructions, use gpt-4o-mini-tts.
OpenAI: TTS-1-HD is a text-to-speech model optimised for high audio fidelity, converting text into natural-sounding spoken audio with greater clarity and detail than the standard tts-1 model, at roughly double the cost. It ships six built-in voices (alloy, echo, fable, onyx, nova, shimmer), supports six output containers (mp3, opus, aac, flac, wav, pcm) and adjustable speech speed from 0.25x to 4.0x. Input is capped at 4096 characters per request. For lower-latency real-time use cases, use tts-1; for voice-style control via natural-language instructions, use gpt-4o-mini-tts.
Whisper-1 is OpenAI’s speech-to-text (and translation) model for audio transcription. It accepts common audio formats and returns text (and optional segments/timestamps). It is used via the Audio Transcriptions/Translations endpoints.
Whisper Large V3 is OpenAI's general-purpose speech recognition model, trained on a large and diverse audio dataset. It is a multi-task model capable of multilingual speech recognition, speech translation to English, and language identification, and returns segment-level and word-level timestamps alongside the full transcription.
Model scores sourced from ArtificialAnalysis