Gemini 1.5 Pro
Google DeepMind · United States · 2024
The model that turned context length into a headline feature — a million tokens, then two, while rivals were still at 128,000.
Gemini 1.5 Pro, announced on 15 February 2024, is the model that made long context a competitive battleground. Where the rest of the industry had settled around 128,000 tokens, Google shipped a standard window of 128,000 with a preview reaching one million — raised to two million at Google I/O in May 2024. A million tokens is roughly an hour of video, eleven hours of audio, thirty thousand lines of code or a very long novel, all held in a single prompt without retrieval tricks. Google also published what many labs would not: the model is a sparse mixture-of-experts transformer, which is how it delivered quality comparable to the much larger Gemini 1.0 Ultra using significantly less compute. In needle-in-a-haystack style retrieval tests it recovered specific facts from these enormous inputs with high reliability, and the technical report described it learning to translate Kalamang — a language with fewer than 200 speakers and almost no presence in training data — from a grammar manual placed in the prompt. It was natively multimodal in a practical sense: text, images, audio and video could be mixed in the same request, which made it the first flagship model that was genuinely useful for analysing recordings and footage rather than stills. The line was followed by Gemini 1.5 Flash, a faster and cheaper distilled variant, and superseded by the Gemini 2.0 family in late 2024. Google began withdrawing Gemini 1.5 Pro from new API projects in April 2025. Its lasting effect is visible in the fact that a large context window is now a standard expectation rather than a differentiator.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!