Google has launched its flagship large language model (LLM) and GPT-4 competitor, Gemini.
Gemini, which was first announced at Google I/O in June, is now generally available to the public and is intended long-term to be integrated across virtually every Google product. Google is stressing Gemini's "multimodal" qualities, which means it can process and leverage different versions of data — not just text, which the average generative AI user will be most familiar with to date, but also images, code, audio and video.
Demis Hassabis, CEO and Co-Founder of Google DeepMind, said in a blog post celebrating the launch:
Gemini is the result of large-scale collaborative efforts by teams across Google, including our colleagues at Google Research. It was built from the ground up to be multimodal, which means it can generalize and seamlessly understand, operate across and combine different types of information including text, code, audio, image and video."
Reports last month suggested that Gemini had been delayed until Q1 2024, so Gemini's launch during its initially planned December date is something of a surprise.
Google has also optimized Gemini in three sizes — Ultra, Pro and Nano, which the tech giant says enables flexibility across use cases, meaning it is "able to efficiently run on everything from data centers to mobile devices". Ultra is Google's largest and most capable model for highly complex tasks, Pro is its most appropriate model for scaling across a wide range of tasks, and Nano is the model best for on-device tasks.
Google also stressed that its Ultra Gemini version surpasses "current state-of-the-art results on 30 of the 32 widely-used academic benchmarks" used in LLM research and development.
"Introducing Gemini 1.0, our most capable and general AI model yet," added Google CEO Sundar Pichai on X. "Built natively to be multimodal, it’s the first step in our Gemini-era of models. Gemini is optimized in three sizes - Ultra, Pro, and Nano. Gemini Ultra’s performance exceeds current state-of-the-art results on 30 of the 32 widely-used academic benchmarks."
Introducing Gemini 1.0, our most capable and general AI model yet. Built natively to be multimodal, it’s the first step in our Gemini-era of models. Gemini is optimized in three sizes - Ultra, Pro, and Nano
Gemini Ultra’s performance exceeds current state-of-the-art results on… pic.twitter.com/pzIw6iCPPN
— Sundar Pichai (@sundarpichai) December 6, 2023
Additionally, Google says that Gemini Ultra is the first LLM to outperform human experts on massive multitask language understanding (MMLU). This framework uses a combination of 57 subjects, including maths, physics, history, law, medicine and ethics for benchmarking knowledge and problem-solving capabilities.
Not missing a trick, Google's announcement blog compares Gemini's MMLU (and other metrics) against OpenAI's GPT-4, with its 90.0 percent MMLU beating GPT-4's 86.4 percent.
Gemini 1.0 is now rolling out across a range of Google products and platforms, including Bard and Google's Pixel 8 Pro device.




