Retrospective Analysis of Google Gemini
摘要
The group of multimodal models Gemini holds the capability to understand image, text, video, and audio files. Gemini comes in many forms, i.e., Nano, Ultra, and Pro sizes for a range of applications starting from complicated logical tasks to use cases of on-device memory constraints. Gemini Ultra model evaluation reflects the benchmark score of 30 of 32 which makes it the initial model to obtain human-level performance on MMLU. This research paper showcases an outline of the model architectonic, training dataset, and training infrastructure. Later detailed results are presented based on the investigation of Gemini model groups, outlining the study for benchmarks and human liking evaluation, i.e., code, text, audio, image, and video which consist of performance in English and multilingual capability. We outline the responsible deployment approach and procedure of impact assessment, development of model policy, investigations, and reduction of harm before deployment decision.