GPT-4o mini Search Preview
OpenAI · United States · 2025
The cheap way to ask the web a question — sixteen times less per token than its full-size twin.
GPT-4o mini Search Preview launched on 11 March 2025 in the same announcement as its full-size twin, and answered the obvious objection to search models: that grounding every answer in live results would be too expensive to run at volume. It cost 15 cents per million input tokens and 60 cents per million output — a sixteenth of GPT-4o Search Preview's rate on input — while keeping the same 128,000-token context window and 16,384-token output ceiling. What it did not do was make search itself cheaper. The per-call fee for each web query was billed separately from tokens, so on short answers, where the tool fee dominates the bill, the saving was far smaller than the headline ratio suggests. That arithmetic — cheap model, fixed-price tool — shaped how developers used it: high-volume, short-answer lookups rather than long research sessions. Input was text, output text, with an October 2023 knowledge cutoff. Like the rest of the March 2025 preview cohort it ran in Chat Completions only, was deprecated on 22 April 2026 and shut down on 23 July 2026, replaced by the built-in web search tool available to any model in the Responses API.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!