{"slug":"google-deepmind-sq","name":"Google DeepMind","logo_url":"https://summerofcode.withgoogle.com/media/org/google-deepmind-sq/cmdhldmexaj4kpms-360.png","website_url":"https://github.com/google-deepmind","tagline":"Google DeepMind's open-source projects","contact_links":[{"name":"email","value":"webpaige@google.com"},{"name":"twitter","value":"https://twitter.com/googleaidevs"}],"date_created":"2025-02-11T17:43:46.445347Z","tech_tags":["python","javascript","typescript","Jax","Gemma"],"topic_tags":["python","Jax","AI,","Gemma"],"categories":["Artificial Intelligence"],"program_slug":"2025","logo_bg_color":null,"description_html":"Google DeepMind is a leading AI research organization committed to solving intelligence and using it to advance science and benefit humanity. We develop cutting-edge machine learning models and techniques, pushing the boundaries of AI across various domains. Our open-source projects, such as JAX (for high-performance numerical computing and machine learning) and Gemma (our open family of models), empower the wider research community. We are dedicated to open science, collaboration, and fostering the next generation of AI talent.","ideas_list_url":"https://goo.gle/deepmind-gsoc-projects-2025","projects":[{"title":"Unified Gemini Example Cookbook: Migrating and Modernizing Open-Source Learning Resources","project_code_url":"https://gist.github.com/andycandy/23856793dafd42e6d3c25b2c999f5217","date_created":"2025-05-06T18:00:50.285385Z","tech_tags":["python","javascript","typescript","Gemini SDK"],"topic_tags":["documentation","tutorial","SDK migration"],"status":"passed","program_slug":"2025","contributor_display_name":"andycandy","mentor_names":["Google DeepMind"],"abstract_short":"This proposal aims to upgrade and expand existing open-source tutorials and examples to support the new unified Gemini SDKs for JavaScript/TypeScript...","abstract_html":"This proposal aims to upgrade and expand existing open-source tutorials and examples to support the new unified Gemini SDKs for JavaScript/TypeScript and Python. My approach involves three key components:\n(i) Cookbook Migration:\nMigrate existing Python examples in the Gemini Cookbook to idiomatic JavaScript/TypeScript. I will ensure the code follows best practices in both languages and is thoroughly documented like the already existing cookbook examples.\n(ii) New Tutorials & End-to-End Examples:\nDevelop fresh, comprehensive tutorials demonstrating a variety of Gemini use cases. Examples include building a chatbot, creating a text summarization tool, and generating images from text prompts. Each tutorial will include step-by-step instructions, code samples, and in-depth explanations of underlying Gemini concepts.\n\nDeliverables include:\n- A set of migrated and newly developed code examples in both Python and JavaScript/TypeScript.\n- Comprehensive, user-friendly tutorials and documentation that clearly explain the code and underlying concepts.\n- Updated examples for existing open-source libraries utilizing the Gemini SDK, with clear migration guides.","date_archived":"2025-05-06T18:00:50.285385Z","id":"5FqYnamd","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Facet AI: No-code web platform to democratize small language model fine-tuning","project_code_url":"https://jetchiang.co/blogs/gsoc/","date_created":"2025-05-06T18:00:53.018089Z","tech_tags":["python","gcp","typescript","terraform","pytorch","HuggingFace","FastAPI","Next.js","Unsloth","Trl"],"topic_tags":["web","machine learning","cloud","LLM","Gemma","Post-training"],"status":"passed","program_slug":"2025","contributor_display_name":"Jet Chiang","mentor_names":["Google DeepMind"],"abstract_short":"Fine-tuning LLMs like Gemma requires deep ML expertise and complex setups, slowing down adoption of SLMs in various industries and communities...","abstract_html":"Fine-tuning LLMs like Gemma requires deep ML expertise and complex setups, slowing down adoption of SLMs in various industries and communities despite their rapid advancements. Existing resources like Colab notebooks aren’t scalable for real-world workflows. Facet AI solves this with a no-code web platform that streamlines the entire fine-tuning process—from dataset curation to model deployment. It supports multiple post-training methods, including full fine-tuning, PEFT, SFT, and RFT, with built-in evaluation and export tools. Powered by Google Cloud, Facet AI removes technical hurdles so teams can focus on building innovative text and multimodal applications.","date_archived":"2025-05-06T18:00:53.018089Z","id":"N2kjH8yi","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Gemini API Developer  Workspace in Postman","project_code_url":"https://github.com/jevonmao/GSoC-2025-jevonmao","date_created":"2025-05-06T18:00:53.591020Z","tech_tags":["javascript","json","rest api","GitHub Actions","POSTMan"],"topic_tags":["API Integration","API documentation","Technical Documentation","Workspace automation"],"status":"passed","program_slug":"2025","contributor_display_name":"Jevon Mao","mentor_names":["Google DeepMind"],"abstract_short":"This project proposes the creation of a comprehensive Postman Workspace tailored for Google’s Gemini API suite. It will offer developers a robust,...","abstract_html":"This project proposes the creation of a comprehensive Postman Workspace tailored for Google’s Gemini API suite. It will offer developers a robust, user-friendly hub to explore, integrate, and test Gemini’s capabilities, significantly lowering the barrier to entry for new Gemini developers.","date_archived":"2025-05-06T18:00:53.591020Z","id":"ntJk8lSk","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Develop Gemini Examples in Swift","project_code_url":"https://gist.github.com/YoungHypo/3dbf4b6553fd6c3e4c94e7c204294cbc","date_created":"2025-05-06T18:01:13.820207Z","tech_tags":["swift","ios","Firebase","Gemini API"],"topic_tags":["artificial intelligence","api","mobile"],"status":"passed","program_slug":"2025","contributor_display_name":"Haibo Yang","mentor_names":["Google DeepMind"],"abstract_short":"This project aims to refactor the firebase/quickstart-ios to demonstrate the latest Firebase AI Logic features, including multimodal analysis,...","abstract_html":"This project aims to refactor the firebase/quickstart-ios to demonstrate the latest Firebase AI Logic features, including multimodal analysis, function calling, and grounding context.\n\nThe key contributions are:\n\t1.\tRebuilt Main UI;\n\t2.\tIntegrated ConversationKit;\n\t3.\tSupported Multimodal Analysis;\n\t4.\tModular feature+example design.\n\nThis project will provide complete examples of using Gemini on iOS, making the quickstart not only quick to run but also easy for new developers to understand and reuse—significantly lowering the learning curve.","date_archived":"2025-05-06T18:01:13.820207Z","id":"lba2Eefe","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Enhancing Gemini API Integrations in OSS Agents Tools","project_code_url":"https://gist.github.com/msaadg/3904e2e74661192951618454f1a32f3a","date_created":"2025-05-06T18:01:16.276085Z","tech_tags":["python","Langchain","LlamaIndex","Gemini API"],"topic_tags":["artificial intelligence","Open Source Software","Multimodal","AI Agent"],"status":"passed","program_slug":"2025","contributor_display_name":"Muhammad Saad (msaadg)","mentor_names":["Google DeepMind"],"abstract_short":"The Gemini API, known for its multimodal capabilities and efficiency, is a powerful tool that enables advanced AI agent interactions. However, its...","abstract_html":"The Gemini API, known for its multimodal capabilities and efficiency, is a powerful tool that enables advanced AI agent interactions. However, its integration within open-source software (OSS) tools like LangChain and LlamaIndex is limited by gaps in functionality and documentation. LangChain and LlamaIndex, two frameworks designed to streamline the development of LLM-powered applications, will significantly benefit from enhanced Gemini API features.\n\nThis project aims to enhance Gemini API integrations in these tools by implementing high-priority features that unlock its full potential. Key improvements include dedicated support for structured output, tool calling with Union types, and resolution of the code execution bug in LangChain, as well as context caching, search grounding, and code execution features for LlamaIndex. In addition to these features, I will update Gemini-specific documentation and create practical workflows to make these tools easier for developers to adopt.\n\nBy addressing these challenges and engaging with the OSS community for feedback, this project will increase Gemini API adoption and empower developers to create powerful, scalable AI agents. The final deliverables include six new features, updated documentation, example workflows, and a community engagement report.","date_archived":"2025-05-06T18:01:16.276085Z","id":"GZnZhIEh","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Streamline experiment execution and improve report UI for OSS-Fuzz-Gen","project_code_url":"https://github.com/google/oss-fuzz-gen","date_created":"2025-05-06T18:02:06.210356Z","tech_tags":["python","javascript","ci","jinja"],"topic_tags":["ai","fuzzing","LLMs"],"status":"passed","program_slug":"2025","contributor_display_name":"Myan (My Anh) Vu","mentor_names":["Dongge Liu"],"abstract_short":"OSS-Fuzz-Gen, a framework using LLMs for fuzz target generation and evaluation by Google, currently has a basic experiment report UI alongside a...","abstract_html":"OSS-Fuzz-Gen, a framework using LLMs for fuzz target generation and evaluation by Google, currently has a basic experiment report UI alongside a manual CI workflow requiring significant oversight. This proposal introduces an improved report interface amongst other features to improve usability and data presentation. Concurrently, the project will streamline the experiment execution workflow for developers and security researchers by introducing more automation into the CI pipeline.","date_archived":"2025-05-06T18:02:06.210356Z","id":"USNSh5o9","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"ATIA: A BENCHMARK FOR ADVERSARIAL TOOL INFILTRATION IN AGENTS","project_code_url":"https://github.com/mattngyn/atia_benchmark","date_created":"2025-05-06T18:02:44.720651Z","tech_tags":["python","pytorch","HuggingFace Transformers","Langchain","UK AISI Inspect"],"topic_tags":["machine learning","benchmarking","nlp","AI Safety","Adversarial Robustness","Multimodel Agents"],"status":"passed","program_slug":"2025","contributor_display_name":"Matthew Nguyen","mentor_names":["Google DeepMind"],"abstract_short":"As multimodal agents become increasingly integrated into real-world applications, ensuring their safe and reliable tool-use behavior is paramount. We...","abstract_html":"As multimodal agents become increasingly integrated into real-world applications, ensuring their safe and reliable tool-use behavior is paramount. We introduce ATIA, a novel evaluation suite designed to systematically assess the vulnerability of tool-calling agents to adversarial inputs. ATIA focuses on crafting malicious multimodal inputs—combinations of images, text, and other modalities—that covertly manipulate an agent's decision-making process regarding external tool invocations. By simulating a variety of attack vectors, including visual prompt injections, image–text conflicts, and spurious cross-modal cues, ATIA quantifies critical metrics such as tool call correctness, unsafe tool invocation rate, and adversarial success rate. This benchmark not only exposes potential weaknesses in current multimodal agent architectures but also provides actionable insights for developing robust, safe, and aligned systems. ATIA aims to serve as a reproducible and extensible framework, driving forward research in adversarial robustness and security for next-generation AI agents.","date_archived":"2025-05-06T18:02:44.720651Z","id":"DcszoT5l","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Xarray-JAX Integration Library","project_code_url":"https://medium.com/@msa242/google-deepmind-gsoc-2025-xarray-jax-c64320e5c1e3","date_created":"2025-05-06T18:02:45.234694Z","tech_tags":["python","numpy","xarray","Jax"],"topic_tags":["machine learning","data science","deep learning","scientific computing","tooling","data structure","Differentiable programming","xarray"],"status":"passed","program_slug":"2025","contributor_display_name":"Mikhail Sinitcyn","mentor_names":["Google DeepMind"],"abstract_short":"This project aims to develop a Python library to support Xarray data (labeled multi-dimensional array library supported by Deepmind) with JAX...","abstract_html":"This project aims to develop a Python library to support Xarray data (labeled multi-dimensional array library supported by Deepmind) with JAX (Google's differentiable programming framework). This will address the lack of support for labeled multi-dimensional data in machine learning frameworks, extending scientific computing capabilities for domains like weather forecasting and financial modeling.\n\nBuilding on Google Deepmind's Xarray-JAX implementation in Graphcast, I will create a standalone library that leverages JAX's recent Python Array API standard implementation to replace the current wrapper-based approach with native integration. The solution will include direct JAX array usage through the API, updated coordinate handling, adapted tree utility registrations, and preservation of existing functionality.\n\nDeliverables include a production-ready xarray-jax library on PyPI, comprehensive test suite, tutorial notebook demonstrating usage, and documentation covering core functionality. This integration will enable scientists and researchers to work with labeled multi-dimensional data in JAX.","date_archived":"2025-05-06T18:02:45.234694Z","id":"PvgGDxQh","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Batch Prediction Framework: Long Context and Context Caching for Video Analysis","project_code_url":"https://github.com/seanbrar/gemini-batch-prediction","date_created":"2025-05-06T18:02:57.616195Z","tech_tags":["python","json","API Design","asynchronous programming","cache management","Gemini API"],"topic_tags":["machine learning","natural language processing","Video Analysis","Context Management","API Optimization","Batch Processing"],"status":"passed","program_slug":"2025","contributor_display_name":"Sean Brar","mentor_names":["Google DeepMind"],"abstract_short":"This project develops an efficient framework for analyzing educational video content using the Gemini API. The approach combines optimized batch...","abstract_html":"This project develops an efficient framework for analyzing educational video content using the Gemini API. The approach combines optimized batch prediction, intelligent context caching, and conversational memory to reduce API usage significantly while improving coherence in multimodal interactions. The framework addresses common challenges when analyzing video content: redundant API calls, context fragmentation with long transcripts, and lack of conversational coherence across related questions. Key deliverables include a modular architecture implementing efficient batch processing with dependency analysis, adaptive context caching strategies for transcripts of varying lengths, conversation memory management for follow-up questions, and structured response formatting with timestamp references. The solution aims to achieve a 4-5× reduction in API calls while maintaining high-quality responses, serving as both a practical tool and a reference implementation demonstrating Gemini API best practices.","date_archived":"2025-05-06T18:02:57.616195Z","id":"dv9GeGEi","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Multimodal Intelligence: Supercharging Agents with Gemini","project_code_url":"https://github.com/Adewale-1/Context_reference_store","date_created":"2025-05-06T18:03:16.671731Z","tech_tags":["JSON Schema","Python, REST APIs, AsyncIO,Caching, Embedding Models, PyTorch,Websockets"],"topic_tags":["Machine Learning, Large Language Models, Multimodal AI, Agent Frameworks","API Orchestration, Function Calling, Natural Language Processing, Context Handling"],"status":"passed","program_slug":"2025","contributor_display_name":"Wale","mentor_names":["Google DeepMind"],"abstract_short":"This project addresses critical gaps in Gemini API integration across leading agent frameworks (LangChain, LlamaIndex, CrewAI, Composio), where...","abstract_html":"This project addresses critical gaps in Gemini API integration across leading agent frameworks (LangChain, LlamaIndex, CrewAI, Composio), where multimodal capabilities and function calling remain underdeveloped or inconsistent. By implementing standardized components for multimodal processing, function calling, and performance optimization, the project will democratize access to Gemini's advanced capabilities for developers building sophisticated AI agents.\n\nThe implementation will follow a systematic approach, beginning with a comprehensive framework audit and gap analysis, followed by development of core multimodal support, function calling capabilities, and performance optimization layers. The solution will include a unified adapter layer that standardizes Gemini API access while respecting each framework's architectural patterns.\n\nDeliverables include: (1) enhanced framework integrations with full multimodal support, (2) standardized function calling implementations, (3) performance optimization components including token management and caching systems, (4) comprehensive documentation and examples, and (5) a benchmarking suite for performance analysis. This project will establish a new standard for Gemini integration in the agent ecosystem, enabling entirely new classes of multimodal AI applications.","date_archived":"2025-05-06T18:03:16.671731Z","id":"C9emc20S","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Enhancing Gemini Integration in Roo Code","project_code_url":"https://gist.github.com/HahaBill/a3dbb27c3b5ec762db0d4580b921e04f","date_created":"2025-05-06T18:03:26.955973Z","tech_tags":["javascript","nodejs","typescript","npm","VSCode Extension API","Gemini SDK"],"topic_tags":["web","machine learning","ui/ux","LLM"],"status":"passed","program_slug":"2025","contributor_display_name":"Ton Hoang Nguyen (Bill)","mentor_names":["Google DeepMind"],"abstract_short":"This project improved Gemini AI integration in Roo Code VS Code extension, serving 800,000+ developers. Key deliverables included real-time Google...","abstract_html":"This project improved Gemini AI integration in Roo Code VS Code extension, serving 800,000+ developers. Key deliverables included real-time Google Search grounding for live web research, progressive model migration eliminating configuration errors, and explicit context caching design reducing token costs by 75%. Additionally, fixed critical streaming race conditions in citation handling and designed cost-effective solutions making AI coding assistants accessible to individual developers and small teams.","date_archived":"2025-05-06T18:03:26.955973Z","id":"XxTqA1r9","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Open-Source Multimodal Benchmarks and Adversarial Robustness Testing for Gemini 2.X models","project_code_url":"https://github.com/sar1kumar/adv-mml-benchmark-suite","date_created":"2025-05-06T18:03:36.540380Z","tech_tags":["python","Jax","Gemini 2.0"],"topic_tags":["Adversarial Attacks","Inference Time Compute","Multi Modal LLM Benchmarking"],"status":"passed","program_slug":"2025","contributor_display_name":"Saravan_Kumar","mentor_names":["Google DeepMind"],"abstract_short":"This project aims to advance the evaluation framework for Google’s Gemini 2.0 and Gemma 3-27B multimodal models by integrating a diverse set of...","abstract_html":"This project aims to advance the evaluation framework for Google’s Gemini 2.0 and Gemma 3-27B multimodal models by integrating a diverse set of open-source, domain-specific benchmarks spanning healthcare, robotics, and general-purpose multimodal reasoning.\nIn the healthcare domain, the evaluation will leverage datasets such as VQA-RAD and OmniMedVQA to assess the models' capabilities in interpreting and reasoning over complex clinical imaging and textual data. For the robotics domain, benchmarks like EmbodiedBench and EmbodiedQA will evaluate the models' proficiency in understanding spatial, visual, and language cues for grounded human-robot interaction tasks. Additionally, general multimodal benchmarks such as SME and BenchLMM will test the models’ ability to produce accurate, context-aware, and human-like explanations grounded in visual inputs.\nTo ensure robustness, the project will also incorporate adversarial attack benchmarks across modalities, evaluating how resilient Gemini 2.0 and Gemma 3-27B are against input perturbations—including visual occlusions, textual prompt manipulations, and conflicting cross-modal signals. Alongside task accuracy, we will analyze the relationship between inference time and adversarial attack success, quantifying how model latency and confidence shift under adversarial conditions.","date_archived":"2025-05-06T18:03:36.540380Z","id":"8r8SJScX","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Evaluate Gemini on an Open-Source Benchmark: OpenUI Eval","project_code_url":"https://github.com/anxkhn/openui_eval_report/blob/main/README.md","date_created":"2025-05-06T18:03:52.237409Z","tech_tags":["python","selenium","git","Puppeteer","ollama","Gemini SDK","HF Transformers"],"topic_tags":["machine learning","benchmarking","generative AI","LLM","Multimodal AI","AI Evaluation"],"status":"passed","program_slug":"2025","contributor_display_name":"@anxkhn (Anas Khan)","mentor_names":["Google DeepMind"],"abstract_short":"This project evaluates large language models on frontend coding and UI generation tasks by extending and integrating ideas from open-source coding...","abstract_html":"This project evaluates large language models on frontend coding and UI generation tasks by extending and integrating ideas from open-source coding benchmarks such as Multi-SWE-bench and WebDev Arena. The evaluation framework assesses coding and UI generation capabilities across various dimensions, with particular emphasis on multimodal features for frontend development. It leverages the models’ ability to generate, render, and iteratively refine solutions using visual feedback from rendered UIs, and supports multiple frontend frameworks (React, Vue, Angular, Svelte, Next.js) alongside single-file HTML tasks.\n\nThe solution involves building a modular pipeline with typed configuration, provider adapters (Ollama, vLLM, OpenRouter), and a model manager; implementing multimodal evaluation loops that incorporate screenshots and structured judging; supporting iterative evaluation (allowing models to refine solutions over multiple attempts); expanding the task taxonomy with increasingly complex frontend challenges (calculators, dashboards, interactive apps, framework-based projects); and automating the entire evaluation pipeline with reproducible runs and detailed artifact logging.\n\nDeliverables include the enhanced evaluation framework with new benchmark tasks, structured outputs and reports (JSON, screenshots, summaries), automation scripts for reproducible runs, and thorough documentation. The framework provides a transparent and extensible way to compare advanced AI coding capabilities in frontend development without relying on a leaderboard, focusing instead on rigorous, reproducible benchmarking.","date_archived":"2025-05-06T18:03:52.237409Z","id":"oZbPWaob","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Exploring & Extending Function Calling in Gemma","project_code_url":"https://gamemaker1.github.io/projects/offline-function-calling/","date_created":"2025-05-06T18:03:59.256412Z","tech_tags":["python","typescript","LLMs"],"topic_tags":["API Design","MCP","LLM Tools","LLM Function Calls","A2A"],"status":"passed","program_slug":"2025","contributor_display_name":"Vedant Kulkarni","mentor_names":["Google DeepMind"],"abstract_short":"This project aims to: (a) investigate to explore technical possibilities, enhance specifications, and find applications for specific use cases and...","abstract_html":"This project aims to: (a) investigate to explore technical possibilities, enhance specifications, and find applications for specific use cases and domains, (b) create a collaborative open source site for Gemma prototypes and cookbooks for the previously identified use cases across domains, and (c) document learnings and create RFCs based on the investigations to provide a roadmap for future work.","date_archived":"2025-05-06T18:03:59.256412Z","id":"rexKK7eu","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"gemini-batcher: A Python package and learning resources for efficient API calls with Gemini LLMs","project_code_url":"https://github.com/phil-daniel/gemini-batcher/blob/main/summary.md","date_created":"2025-05-06T18:03:59.749082Z","tech_tags":["python","Gemini"],"topic_tags":["Large Language Models","Code samples","Context Caching","Batching"],"status":"passed","program_slug":"2025","contributor_display_name":"Phillip Daniel","mentor_names":["Google DeepMind"],"abstract_short":"The aim of this project is to develop a comprehensive learning resource for developers working with the Gemini Python SDK. This would consist of a...","abstract_html":"The aim of this project is to develop a comprehensive learning resource for developers working with the Gemini Python SDK. This would consist of a range of documentation, interactive code samples and an easy to use pre-written package that demonstrate cover a range of topics which can be used to efficiently manage long content, one of the most common challenges when building with large language models. It will provide an overview of techniques such as batching, chunking and handling various media types (such as text, audio and video) in addition to highlighting best practices for techniques such as caching and error handling.","date_archived":"2025-05-06T18:03:59.749082Z","id":"QwrjzUxs","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"VS Code Extension to assist with coding powered by Gemini","project_code_url":"https://github.com/krishnaagrawal7508/GeminiBot","date_created":"2025-05-06T18:04:41.871084Z","tech_tags":["javascript","react","typescript","node","apis","JEST","VS Code","Lru-cache"],"topic_tags":["artificial intelligence","ai","chat","Developer Productivity","Extension","gemini","VS code","Coding Assistance"],"status":"passed","program_slug":"2025","contributor_display_name":"krishnaagrawal","mentor_names":["Google DeepMind"],"abstract_short":"A VS Code (or JetBrains) extension to provide AI-powered coding assistance using Google’s Gemini API. This tool enhances the developer experience by...","abstract_html":"A VS Code (or JetBrains) extension to provide AI-powered coding assistance using Google’s Gemini API. This tool enhances the developer experience by offering smart code completions, real-time debugging help, and natural language to code conversion. It can suggest optimized code, detect errors, explain complex logic, and help developers write better, more efficient programs. The extension will also improve response times with caching mechanisms and provide a seamless UI/UX for smooth interaction. By integrating AI-driven support directly into VS Code, this project aims to boost productivity, reduce debugging time, and make coding more intuitive and accessible for developers.\nHere is the source code of the extension created by me, in the GitHub repository here: https://github.com/krishnaagrawal7508/GeminiBot","date_archived":"2025-05-06T18:04:41.871084Z","id":"f9mtEUFy","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Improving Gemini Documentation for Open Source Model Providers Promptfoo and Weights & Biases","project_code_url":"https://adel-muursepp.medium.com/google-summer-of-code-2025-with-google-deepmind-promptfoo-47f48917af06","date_created":"2025-05-06T18:04:53.706231Z","tech_tags":["python","javascript","github","LLMs","FULL STACK DEVELOPMENT","Gemini","Technical documentation"],"topic_tags":["Technical Documentation","LLM Evalution","LLM Safety"],"status":"passed","program_slug":"2025","contributor_display_name":"Adel Muursepp","mentor_names":["Google DeepMind"],"abstract_short":"The project will close the documentation and evaluation gap for Google’s Gemini models by contributing structured onboarding guides and benchmarking...","abstract_html":"The project will close the documentation and evaluation gap for Google’s Gemini models by contributing structured onboarding guides and benchmarking templates to Promptfoo and Weights & Biases Weave, with specifically introducing models like Gemini 2.5 Pro, Gemini 2.0 Flash and Gemini 2.0 Flash-Lite. Despite Gemini’s powerful features—like multimodality and advanced safety settings—its presence in open-source evaluation tools lags behind GPT or Llama models. By improving usability, comparability, and safety transparency, this work will help developers, researchers, and product teams integrate and assess Gemini effectively, making it a fully accessible option in the LLM ecosystem.","date_archived":"2025-05-06T18:04:53.706231Z","id":"TwyQFyTr","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"SciResearchBench: A Multimodal Benchmark for Scientific Reasoning and Discovery","project_code_url":"https://alampara.com/gsoc2025/","date_created":"2025-05-06T18:05:17.818221Z","tech_tags":["python"],"topic_tags":["Multi-modal evaluation, LLMs , Foundation Models, Scientific Discovery, Reasoning, Benchmark"],"status":"passed","program_slug":"2025","contributor_display_name":"Nawaf Alampara","mentor_names":["Google DeepMind"],"abstract_short":"Scientific discovery fundamentally relies on integrating and reasoning over multimodal information—text, diagrams, plots, spectra, microscopy images,...","abstract_html":"Scientific discovery fundamentally relies on integrating and reasoning over multimodal information—text, diagrams, plots, spectra, microscopy images, and experimental observations. While multimodal Language Models like Gemini show promise as potential AI research assistants, their capabilities in nuanced scientific reasoning remain largely unevaluated, particularly across diverse scientific domains. Existing benchmarks often focus on general knowledge, single modalities, or lack the depth needed to probe scientific reasoning. This project proposes the creation of SciResearchBench, a novel,\nopen-source multimodal benchmark specifically designed to evaluate the scientific reasoning and discovery capabilities of multimodal across key research domains like Chemistry, Materials Science, Biology, and Physics. SciResearchBench will mirror real-world research workflows, encompassing tasks in multimodal data integration, experimental understanding, results interpretation, and hypothesis generation. A core focus will be extensive ablation studies to systematically analyze model sensitivities (prompting, cross-modal integration, grounding, reasoning complexity, context length),\nproviding crucial insights into model failure modes and avenues for improvement. The project deliverables include the curated/generated dataset, robust evaluation code, comprehensive evaluation results for Gemini models, detailed analysis, and the full open-source release of all components to the research community.","date_archived":"2025-05-06T18:05:17.818221Z","id":"xshu9ha6","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Batch Prediction with Long Context and Context Caching","project_code_url":"https://gist.github.com/vanshksingh/0f4db89bfb8fc96adc90d8e779448fd2","date_created":"2025-05-06T18:05:21.695091Z","tech_tags":["python","json","markdown","Streamlit","asyncio","Caching","FAISS","RAG","Google Gemini API"],"topic_tags":["education","machine learning","developer tools","cloud","LLM","Gen-AI","AI Tooling"],"status":"passed","program_slug":"2025","contributor_display_name":"vanshksingh","mentor_names":["Google DeepMind"],"abstract_short":"This project developed an open-source code sample for batch question answering on long-form transcripts (such as lectures or documentaries) using...","abstract_html":"This project developed an open-source code sample for batch question answering on long-form transcripts (such as lectures or documentaries) using Google Gemini APIs. It delivers a modular pipeline for batch predictions with asynchronous processing, long-context handling through transcript chunking and fallback strategies, a context caching system to reduce token usage, support for interconnected and multi-turn questions, and structured Markdown/JSON output with timestamp references. Several stretch goals were also completed, including Streamlit user interfaces for batch runs and cache management, efficiency benchmarking with reproducible token savings, and cache-aware batching strategies. During testing, the project surfaced two important API issues: an explicit caching bug on free-tier accounts (escalated internally as a P0 issue) and the lack of documentation around batch API availability on free-tier accounts. This work provides a practical reference for developers building educational tools, video summarization assistants, and retrieval-augmented generation (RAG) systems with Gemini.","date_archived":"2025-05-06T18:05:21.695091Z","id":"NxQQgsTG","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Crisis Response Toolkit for Gemma Models","project_code_url":"https://gist.github.com/rorosaga/a558aafcdb59aa1fdc7e02362d694308","date_created":"2025-05-06T18:05:35.203356Z","tech_tags":["python","javascript","react","html/css","NextJs","FastAPI","REST APIs"],"topic_tags":["machine learning","web development","ai","healthcare","Crisis Response","Function Calls","tool kit"],"status":"passed","program_slug":"2025","contributor_display_name":"Rodrigo Sagastegui","mentor_names":["Google DeepMind"],"abstract_short":"This project explores how lightweight Gemma models can be applied in life-critical situations through an open-source Crisis Response Toolkit. The...","abstract_html":"This project explores how lightweight Gemma models can be applied in life-critical situations through an open-source Crisis Response Toolkit. The main deliverable, One Minute Agent, is an offline AI assistant designed to provide first-aid guidance during emergencies, bridging the gap before responders arrive and helping improve survival rates. The work also investigates offline deployment, function-calling abilities, and agent workflows to demonstrate how Gemma can support real-time decision-making in crisis scenarios.","date_archived":"2025-05-06T18:05:35.203356Z","id":"sfMAOShm","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Reproducibility as Accuracy (RaA) Benchmark","project_code_url":"https://gist.github.com/pranavagrawaI/dbd7f57adc7316edb5a1d1c2862867f4","date_created":"2025-05-06T18:05:47.569406Z","tech_tags":["python","git","jupyter","Transformers","Python Libraries"],"topic_tags":["machine learning","computer vision","natural language processing","Multimodal AI","Evaluation Metrics"],"status":"passed","program_slug":"2025","contributor_display_name":"Pranav Agrawal","mentor_names":["Google DeepMind"],"abstract_short":"Reproducibility as Accuracy (RaA) is a benchmark which aims to evaluate how effective multimodal AI systems are in preserving information fidelity...","abstract_html":"Reproducibility as Accuracy (RaA) is a benchmark which aims to evaluate how effective multimodal AI systems are in preserving information fidelity during iterative transformations from image to text.  By implementing a recursive process—converting an image to text and back to an image over multiple iterations—the benchmark will apply qualitative and quantitative metrics to assess semantic drift and information degradation. Deliverables include an open-source benchmarking framework, a suite of evaluation metrics, documentation, and analytical reports, providing insights to enhance the reliability of multimodal AI systems.","date_archived":"2025-05-06T18:05:47.569406Z","id":"hoyVl52a","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Creation of a Creative Thinking Benchmark","project_code_url":"https://github.com/theGreen-Coder/MCTB/blob/main/GSoC25.md","date_created":"2025-05-06T18:05:49.454468Z","tech_tags":["python","pytorch","Hugging Face"],"topic_tags":["LLMs","LLM Benchmark Evalution","LLM Creative Thinking"],"status":"passed","program_slug":"2025","contributor_display_name":"Green Code","mentor_names":["Google DeepMind"],"abstract_short":"The goal of this project is to develop a multi-modal and open-source benchmark with which to evaluate Gemini 2.0. Open-source benchmarks are an...","abstract_html":"The goal of this project is to develop a multi-modal and open-source benchmark with which to evaluate Gemini 2.0. Open-source benchmarks are an unbiased way to test the ability of LLMs in different modalities. Current LLM benchmarks are mostly based on reasoning tasks and maths. However, there does not exist a gold standard LLM benchmark to assess creative thinking. This is a critical gap since creative thinking is essential for progress toward artificial general intelligence (AGI). To solve this, the project aims to extend current creative thinking LLM benchmarks and create a gold standard benchmark across four different modalities (text, image, video, and audio).\n\nUpon the project’s completion, a thorough multi-modal open-source benchmark for creative thinking and creative divergence will be created. This will be accompanied by the evaluation metrics of different models (including Gemini 2.0), along with documentation and custom evaluation scripts. Finally, an educational video explanation of the benchmark will be shared on my YouTube channel (https://www.youtube.com/@Green-Code/), which will serve as an introduction to developers and newcomers to the LLM field.","date_archived":"2025-05-06T18:05:49.454468Z","id":"TSDUDJNY","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Self-Contained OSS-Fuzz Module for Researchers","project_code_url":"https://docs.google.com/document/d/12NxFpmMO9Qat-f_5V4hZ-q0v8a8BL9xN9HBMIzJlgQA/edit?usp=sharing","date_created":"2025-05-06T18:06:01.439668Z","tech_tags":["python","PyPI","Google Cloud Storage"],"topic_tags":["security","fuzzing","software development","API Design","Software Refactoring"],"status":"passed","program_slug":"2025","contributor_display_name":"Zewei Wang","mentor_names":["Dongge Liu"],"abstract_short":"This project aims to develop a standalone Python SDK that provides researchers with a streamlined and well-documented API for interacting with...","abstract_html":"This project aims to develop a standalone Python SDK that provides researchers with a streamlined and well-documented API for interacting with OSS-Fuzz. The current access methods to OSS-Fuzz's features and data are fragmented and inconsistent, posing challenges in terms of code duplication, error handling, and usability for new researchers. By encapsulating OSS-Fuzz interactions within a cohesive module, this project will standardize data formats, simplify API usage, and improve accessibility for researchers.","date_archived":"2025-05-06T18:06:01.439668Z","id":"51tdKuLz","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Open-source Gemini Example Apps","project_code_url":"https://gist.github.com/FallenDeity/be473dfbfe75d76483ec086a5765c0ca","date_created":"2025-05-06T18:06:31.624288Z","tech_tags":["python","javascript","github","git","typescript","Jupyter Notebooks","Gemini SDKs"],"topic_tags":["machine learning","computer vision","documentation","Technical Writing","API Showcasing"],"status":"passed","program_slug":"2025","contributor_display_name":"Triyan Mukherjee","mentor_names":["Google DeepMind"],"abstract_short":"The Gemini Cookbook is a set of sample applications and tutorials illustrating different functionalities of the Gemini APIs. The intent of this...","abstract_html":"The Gemini Cookbook is a set of sample applications and tutorials illustrating different functionalities of the Gemini APIs. The intent of this proposal is to modernize current tutorials and documentation to assist with the new unified Gemini SDKs for JavaScript/TypeScript and Python. This involves bringing examples from Python into JS/TS, creating new end-to-end tutorials, and modernizing examples for other open source libraries utilizing outdated versions of the Gemini APIs. By giving users new, current learning tools, this project hopes to facilitate adoption of the Gemini APIs.\n\n1. Migrate Existing Python Examples to TypeScript\nConvert the current Python-based examples in the Gemini Cookbook to idiomatic TypeScript using the latest Gemini SDK. This will ensure cross-language support and make sure the resources are accessible to a wider range of developers.\n2. Create End-to-End Tutorials for Real-World Use Cases\nDesign and implement comprehensive, practical tutorials that showcase the capabilities of the Gemini APIs. These tutorials will include projects such as intelligent chatbots, document summarization tools, and image generation from prompts, covering both basic and advanced use cases. New ideas for tutorials can be submitted by the community, and can be sourced from the existing cookbook issues.\n3. Modernize Open-Source Libraries Using Gemini APIs\nIdentify and contribute to popular open-source projects that use deprecated versions of the Gemini SDKs. Update the examples and documentation to leverage the features and improvements introduced in the latest unified Gemini SDKs.","date_archived":"2025-05-06T18:06:31.624288Z","id":"VhcMTva8","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Gemma Scout: On-device AI camping & wildlife survival companion","project_code_url":"https://www.kaggle.com/competitions/google-gemma-3n-hackathon/writeups/gemma-scout","date_created":"2025-05-06T18:07:11.773459Z","tech_tags":["python","github","bash"],"topic_tags":["machine learning","benchmarking","optimization","edge","inference","Large Language Models","Multimodal","evaluation","On-Device"],"status":"passed","program_slug":"2025","contributor_display_name":"Ryan Rong","mentor_names":["Google DeepMind"],"abstract_short":"About 88 million U.S. households now identify as campers, and ~54 million households took a trip last year National parks run thousands of...","abstract_html":"About 88 million U.S. households now identify as campers, and ~54 million households took a trip last year\n\nNational parks run thousands of search-and-rescue operations every year (e.g., ~3,400 in 2022), and research shows tens of thousands of people were lost in parks over the 2004–2014 decade.\n\nFirst aid, navigation, shelter, and water skills reduce risk and help you protect others. That’s why I built Gemma Scout: an on-device, privacy-first AI assistant built for wilderness survival and outdoor adventures. Powered by a finetuned version of Gemma 3N, it combines domain-specific survival knowledge with real-time multimodal reasoning to guide users through critical tasks like plant identification, mushroom safety, shelter building, and first aid.\n\nThe model was trained on over 30 million tokens and 7,618 curated Q&A pairs, sourced from expert handbooks and image datasets of edible plants and mushrooms. Leveraging Unsloth’s optimized finetuning framework with LoRA and quantization, Gemma Scout runs efficiently on mobile devices without internet access.\n\nUsing LLM evaluations, Gemma Scout outperformed its base model in survival QA by a win rate of 89%, and achieved over 36% improvement in mushroom classification and 47% in plant identification. It supports multimodal chat via llama.cpp mtmd and Swift integration, enabling users to send images and receive grounded, expert-level responses—completely serverless.\n\nWhether you're camping, hiking, or lost in the wild, Gemma Scout helps you make smart, safe decisions—even when you're off the grid.","date_archived":"2025-05-06T18:07:11.773459Z","id":"RSjjE3tM","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"HALO: Hierarchical Abstraction for Longform Optimization","project_code_url":"https://github.com/jeet-dekivadia/google-deepmind","date_created":"2025-05-06T18:07:33.394606Z","tech_tags":["python","redis","pytorch","pandas","HuggingFace Transformers","asyncio","Whisper","FAISS","Pydantic","Gemini API","CLIP","pyannote","Reinforcement Learning (PPO)","Multimodal Learning"],"topic_tags":["machine learning","reinforcement learning","video processing","Conversational AI","Asynchronous Programming","API Optimization","Batch Processing","Context Caching","Video Q&A Systems","Long-Context Modeling","Error Resilience","Distributed Caching"],"status":"passed","program_slug":"2025","contributor_display_name":"Jeet Dekivadia","mentor_names":["Google DeepMind"],"abstract_short":"HALO (Hierarchical Abstraction for Longform Optimization) is an MIT-licensed Python package for efficient large-scale video content analysis,...","abstract_html":"HALO (Hierarchical Abstraction for Longform Optimization) is an MIT-licensed Python package for efficient large-scale video content analysis, installable via pip install halo-video. The system implements a hierarchical processing architecture that reduces computational complexity to O(n log n) using dynamic content-based segmentation. Key components include: a multi-modal fusion pipeline integrating visual (768-dimensional) and transcript (1024-dimensional) embeddings; a three-tier caching system utilizing in-memory hash tables, disk-based serialization, and compressed vector storage; and an optimization layer for API request batching. The architecture maintains content coherence through a sliding context window with 30% chunk overlap. Performance metrics show 93% reduced processing time, 85% lower computational costs, and 98% fewer API calls compared to baseline approaches. HALO supports cross-platform operation (Linux, macOS, Windows), implements error handling with exponential backoff, and provides standardized interfaces for integration with existing ML pipelines. The package processes video content with bounded memory requirements regardless of input length.","date_archived":"2025-05-06T18:07:33.394606Z","id":"5ni7RZ48","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Gemma Facet: End-to-End Fine-Tuning Platform for Gemma Models","project_code_url":"https://www.adarshdubey.com/blog/gsoc-final-report","date_created":"2025-05-06T18:08:32.071882Z","tech_tags":["python","pytorch","Google Cloud Platform","FastAPI","Next.js","Hugging Face","LLMs","Unsloth","Gemma Models"],"topic_tags":["Web Interface","LLMs","Fine tuning","Gemma","SLMs"],"status":"passed","program_slug":"2025","contributor_display_name":"Adarsh J. Dubey","mentor_names":["Google DeepMind"],"abstract_short":"Gemma Facet is a comprehensive platform that provides an end-to-end solution for fine-tuning Gemma language models through a microservices...","abstract_html":"Gemma Facet is a comprehensive platform that provides an end-to-end solution for fine-tuning Gemma language models through a microservices architecture. The platform implements four core services: dataset preprocessing with support for local and Hugging Face datasets, automated training jobs using Unsloth and Transformers libraries, model inference capabilities, and flexible export functionality supporting multiple formats (adapters, merged models, and GGUF).\nThe backend leverages Google Cloud Run services for scalable compute, Firestore for database operations, and Google Cloud Storage for artifact management. The frontend is built with Next.js, Tailwind CSS, and Shadcn UI, providing an intuitive dashboard interface. Infrastructure is managed through Terraform with containerized services using Docker.\nThe system handles complex workflows including dataset splitting configurations, asynchronous training job management, and multi-format model export pipelines. Key technical implementations include Cloud Run Job integration for resource-intensive operations, comprehensive API design with full documentation, and optimized data processing pipelines for efficient model fine-tuning workflows.","date_archived":"2025-05-06T18:08:32.071882Z","id":"6ep1Zcf2","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Develop a Gemini Workspace in Postman","project_code_url":"https://gist.github.com/AniketS02/fce3b866d1b714391857bc6cbf267563","date_created":"2025-05-06T18:08:32.515344Z","tech_tags":["javascript","json","github","POSTMan","Rest API interaction"],"topic_tags":["API Integration","Github action","Postman Workspace"],"status":"passed","program_slug":"2025","contributor_display_name":"Aniket.Saxena","mentor_names":["Google DeepMind"],"abstract_short":"This project aims to create a Gemini Workspace in Postman for interacting with the Gemini API’s and providing a central hub for exploration ,...","abstract_html":"This project aims to create a Gemini Workspace in Postman for interacting with the Gemini API’s and providing a central hub for exploration , integration and troubleshooting. This will reduce the onboarding time for developers to get started with Gemini and the initial learning curve. The key features of this project are : \n\n1. Pre-build collections for every feature and API requests\n2.Examples for every requests to analyse the expected response\n3. Environments for testing and production\n4. Quickstart guide and documentation of every requests - purpose and expected behaviour\n5.Test scripts to check the format of the JSON response and the status code\n6. Mock servers using the postman mocking capabilities\nAutomation via Github Action.\n\nWithout the Gemini Workspace in Postman, developers face manual setup, debugging difficulties, and a steep learning curve. The project simplifies API testing, automation, and collaboration, making the Gemini API more accessible, efficient, and user-friendly","date_archived":"2025-05-06T18:08:32.515344Z","id":"yrmj0F4J","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Gemma Model Fine-tuning UI","project_code_url":"https://github.com/drink970082/GSoC-2025-Gemma-Model-Fine-tuning-UI","date_created":"2025-05-06T18:08:41.797537Z","tech_tags":["python","Streamlit","TensorFlow/PyTorch"],"topic_tags":["machine learning","web development"],"status":"passed","program_slug":"2025","contributor_display_name":"Chen-Hao Wu","mentor_names":["Google DeepMind"],"abstract_short":"Gemma is a lightweight, open-source large language model by Google DeepMind. This project aims to build an intuitive web interface for fine-tuning...","abstract_html":"Gemma is a lightweight, open-source large language model by Google DeepMind. This project aims to build an intuitive web interface for fine-tuning Gemma models. The interface will allow users to upload datasets, configure hyperparameters, monitor training progress, and export trained models — all without writing a single line of code. By lowering the entry barrier, the UI will empower a broader range of users to experiment with and adapt large language models to their specific tasks.","date_archived":"2025-05-06T18:08:41.797537Z","id":"vBF5rvcZ","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"EchoGem – Teaching Gemini to Think in Batches by Prioritizing What Matters","project_code_url":"https://github.com/aryan-410/EchoGem/tree/main","date_created":"2025-05-06T18:08:59.290288Z","tech_tags":["python","flask","tensorflow","celery","numpy","pytorch","nltk","scikit-learn","pandas","FastAPI","Ray","asyncio","spaCy","Qdrant","Langchain","Pinecone","Hugging Face Transformers","FAISS","Gemini","Vector Databases","Gemini API","sentence-transformers","SentenceTransformers","BERTopic","TQDM"],"topic_tags":["machine learning","distributed systems","natural language processing","open source","information retrieval","parallel computing","parallelization","Semantic Search","neural search","Transformer models","question answering","Asynchronous Programming","Large Language Models","Vector similarity","Prompt Engineering","Context Management","Caching Strategies","Topic Modeling","Caching Systems","Batch Processing","Context Caching","Model Efficiency","Token Optimization","Context-Aware AI","Text Embeddings","Batch Inference","Systems Optimization","Pipeline Optimization","Transcript Analysis","Semantic Chunking","Inference Engines","Latency Reduction","Entropy-Based RankingLong-Context Processing"],"status":"passed","program_slug":"2025","contributor_display_name":"Aryan Saboo","mentor_names":["Google DeepMind"],"abstract_short":"EchoGem introduces a novel batching engine designed to answer multiple questions about the same source parallelly to reduce response times heavily....","abstract_html":"EchoGem introduces a novel batching engine designed to answer multiple questions about the same source parallelly to reduce response times heavily. It does so by building a modular smart batching engine for Gemini that teaches the model to think in context-aware batches.\r\n\r\nWhat sets EchoGem apart is its focus on modularity and testing—each component is designed to be independently swappable and improvable allowing for open testing and matching different strategies for different parts of the engine. Instead of a naive sliding window or top-k search approach, EchoGem uses semantic clustering of both chunks and questions to form intelligent batch groups. It is also committed to measurable, reproducible gains. Every design decision—whether it's a new chunking strategy or a different context ranking model—can be rigorously backtested using a structured evaluation suite.","date_archived":"2025-05-06T18:08:59.290288Z","id":"7pu5jVUL","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"AI Evaluations using Gemini APIs","project_code_url":"https://docs.google.com/document/d/1BiBH0Dl1IvYHJomwyXBF_syvHsb_C7kkb7n08-GZkp4/edit?usp=sharing","date_created":"2025-05-06T18:09:13.526705Z","tech_tags":["python","typescript","Gemini"],"topic_tags":["LLM","LLM Evaluation"],"status":"passed","program_slug":"2025","contributor_display_name":"Siddharth Sahu","mentor_names":["Google DeepMind"],"abstract_short":"Current manual evaluation methods for LLM-based AI applications are unsustainable and resource-intensive. While using LLMs as judges offers a...","abstract_html":"Current manual evaluation methods for LLM-based AI applications are unsustainable and resource-intensive. While using LLMs as judges offers a promising alternative, evaluation frameworks provide structured tools that harness LLMs for automated assessments, offering metrics, benchmarks, and scalable testing methodologies that overcome manual evaluation limitations. Gemini's integration with evaluation platforms lacks comprehensive documentation compared to providers like OpenAI and Anthropic. Analysis of ten major frameworks reveals scattered integration, decentralized documentation, and limited implementation guidance for Gemini, creating barriers for developers and researchers.\n\nI propose a three-pronged approach to improve Gemini support across evaluation frameworks. First, adding native Gemini support to select frameworks (DeepEval and TruLens) that already offer direct integrations for other providers. Second, enhancing both abstraction-layer and framework-specific documentation, covering API key management, model configuration, and multimodal capabilities. Third, creating framework-specific guides in Python and TypeScript that demonstrate effective combination of library features with Gemini APIs.\n\nFor code deliverables, I will develop native Gemini integration for DeepEval and TruLens frameworks, with pull requests submitted to respective GitHub repositories. Documentation deliverables include enhanced guidance for abstraction layers (LiteLLM, LangChain, LlamaIndex) focused on Gemini configuration, and framework-specific documentation covering API management, configuration options, multimodal integration, and evaluation best practices.\nTutorial deliverables will include comprehensive guides for all ten analyzed frameworks, with Python implementations for all frameworks and TypeScript for supported ones, step-by-step configuration instructions, advanced usage examples highlighting Gemini's multimodal capabilities, and error handling techniques.","date_archived":"2025-05-06T18:09:13.526705Z","id":"IG6jL3Hj","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Creating The First Benchmark for Evaluating LLMs Across All Five Foundational AI Agent Types","project_code_url":"https://www.namchittai.com/gsoc2025","date_created":"2025-05-06T18:09:18.064909Z","tech_tags":["python","numpy","Large Language Models","vLLM","Multimodal Large Language Models"],"topic_tags":["machine learning","benchmarking","Large Language Models","AI Agents","Multimodal Large Language Models"],"status":"passed","program_slug":"2025","contributor_display_name":"Nattaput Namchittai","mentor_names":["Google DeepMind"],"abstract_short":"With the rise in popularity of AI agents, there is an increasing need for agentic benchmarks. There are 5 foundational types of AI agents that many...","abstract_html":"With the rise in popularity of AI agents, there is an increasing need for agentic benchmarks. There are 5 foundational types of AI agents that many AI agents can be classified as (Simple Reflex, Model-Based Reflex, Goal-Based, Utility-Based, and Learning Agents). Modern agentic workflows are built upon these foundational types or involve hybrids or multi-agent systems of these types. However, there is no existing benchmark that evaluates how LLMs perform generally as each agent type.  Existing AI agent benchmarks are very task-specific, meaning that those wanting to choose the right LLM for agentic tasks which have not yet been benchmarked need to rely on heuristics from similar task-specific data that is available. This novel benchmark aims to evaluate agent performance on various tasks specific to each agent type in order to gauge the generalized performance of LLMs as each specific agent type. Such a benchmark will not only be useful for developers for selecting the right LLM for their agent type but also for researchers to better understand the current agentic capabilities and limitations of existing LLMs. The benchmark will also help us learn which LLMs are more versatile or specialized in agentic settings.","date_archived":"2025-05-06T18:09:18.064909Z","id":"1E6DMVwi","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Creating New Agent Architectures for Concordia","project_code_url":"https://sycorpia.substack.com/p/building-a-multi-agent-negotiation","date_created":"2025-05-06T18:09:41.114311Z","tech_tags":["Python, PyTorch, Concordia, Gemini/Gemma"],"topic_tags":["Agent Architecture, Long Context, Agents, Cooperative AI"],"status":"passed","program_slug":"2025","contributor_display_name":"tesims","mentor_names":["Google DeepMind"],"abstract_short":"The goal of this project is to help strengthen the Concordia framework by developing and open-sourcing a collection of new language model agent...","abstract_html":"The goal of this project is to help strengthen the Concordia framework by developing and open-sourcing a collection of new language model agent architectures. The project is meant to iterate on the work done by contestants during the Concordia Contest at NeurIPS 2024. \r\n\r\nWhile the Concordia Contest 2024 has concluded, the need for well-documented and diverse examples of cooperative agents in the Concordia ecosystem remains important for the ongoing research in cooperative AI. The main benefit being that the new agents created will help lower the barrier of entry for new researchers and engineers, with the hope of inspiring more people to contribute unique approaches to using the framework and speed up the research being done in cooperative AI by providing practical and reusable agent implementations.","date_archived":"2025-05-06T18:09:41.114311Z","id":"j5h3cWWt","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Enhanced Benchmark for Evaluating Intuitive Physics Understanding in Gemma Multimodal Models","project_code_url":null,"date_created":"2025-05-06T18:10:07.037130Z","tech_tags":["python","opencv","git","docker","ffmpeg","Jax","Flax","Video Processing"],"topic_tags":["deep learning","benchmarking","Self-Supervised Learning","Multimodal AI","Intuitive Physics Understanding"],"status":"passed","program_slug":"2025","contributor_display_name":"lucas-maes","mentor_names":["Google DeepMind"],"abstract_short":"This project aims to develop a more rigorous and focused evaluation testbed than that used by Garrido et al. (2025), with the specific goal of...","abstract_html":"This project aims to develop a more rigorous and focused evaluation testbed than that used by Garrido et al. (2025), with the specific goal of assessing the intuitive physics understanding of Gemma models. By addressing key shortcomings in existing evaluations—such as overreliance on textual outputs, limited control over task difficulty, and incomplete coverage of fundamental physical concepts—this new benchmark will enable more accurate and fair comparisons. Ultimately, it will offer deeper insights into what Gemma models actually grasp about the physical world and guide future improvements in their design and training.","date_archived":"2025-05-06T18:10:07.037130Z","id":"g4DnP02P","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Gemma Garage:  Leveraging Gemma 3 to democratize LLM Fine-tuning","project_code_url":"https://gist.github.com/Lucas-Fernandes-Martins/3932c23d9aadeefe267f617318f1e51c","date_created":"2025-05-06T18:18:58.844697Z","tech_tags":["python","javascript","google cloud","react","pytorch","Gemma","PEFT","Transfomers"],"topic_tags":["ai","ui","fullstack","fine-tuning","LLM"],"status":"passed","program_slug":"2025","contributor_display_name":"Lucas Martins","mentor_names":["Google DeepMind"],"abstract_short":"This proposal aims to develop the Gemma LLM Garage, a full-stack interface to manage datasets and fine-tune Gemma models. Its main goal is to...","abstract_html":"This proposal aims to develop the Gemma LLM Garage, a full-stack interface to\nmanage datasets and fine-tune Gemma models. Its main goal is to abstract the technical complexity from the end user, empowering anyone with a computer to fine-tune their models.\n\nSo far, I have implemented basic features such as fine-tuning with LoRA, a mechanism for uploading datasets, simple data augmentation options, a live loss graph, and a simple inference platform to test the fine-tuned model. The main deliverables are implementing data cleaning strategies, improving the data augmentation options, implementing a testing workbench for the fine-tuned models, and supporting alternative fine-tuning methods (e.g. QLoRA).\n\nThe project prototype is currently deployed through Google Cloud and can be accessed here: \nhttps://gemma-garage.web.app/\n\nA quick video demo is available here: \nhttps://drive.google.com/file/d/1Knt9a16NfJTUX5rddte_SHcMDKCzA8NS/view","date_archived":"2025-05-06T18:18:58.844697Z","id":"yT16LTpy","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Enhance Gemini API Integrations in OSS Agents Tools","project_code_url":"https://github.com/andyli11/llamaindex-quickstart","date_created":"2025-05-06T18:19:15.427075Z","tech_tags":["python","Langchain","LlamaIndex"],"topic_tags":["AI Agents"],"status":"passed","program_slug":"2025","contributor_display_name":"Andy L","mentor_names":["Google DeepMind"],"abstract_short":"This project will elevate Gemini API support across widely used open-source agent frameworks like LangChain, LlamaIndex, CrewAI, and...","abstract_html":"This project will elevate Gemini API support across widely used open-source agent frameworks like LangChain, LlamaIndex, CrewAI, and Composio—bridging the gap between Gemini’s full capabilities and the tools developers use to build with it. By introducing support for advanced features such as multimodal inputs, structured function calling, and streaming outputs, I aim to empower developers to build highly interactive, context-aware agents with ease. The project will include contributions to core library integrations, detailed documentation improvements, and the development of end-to-end examples showcasing Gemini in action. The result will be a robust, user-friendly ecosystem for Gemini-based agent development—making it easier than ever for developers to harness its full potential within their own workflows, products, and research.","date_archived":"2025-05-06T18:19:15.427075Z","id":"P4729El9","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Gemma Chat Gradio Demo","project_code_url":"https://huggingface.co/spaces/AC2513/gemma-demo","date_created":"2025-05-06T18:19:33.126411Z","tech_tags":["python","pytorch","Transformers","Pytest","Gradio","Playwright","Hugging Face Spaces"],"topic_tags":["machine learning","ai","chatbot","LLM","Gemma"],"status":"passed","program_slug":"2025","contributor_display_name":"AndyC","mentor_names":["Google DeepMind"],"abstract_short":"A majority of the Gemma chat applications on Hugging Face Spaces do not allow the user to adjust generation settings or system prompts, giving the...","abstract_html":"A majority of the Gemma chat applications on Hugging Face Spaces do not allow the user to adjust generation settings or system prompts, giving the user only a generic experience. This project aims to provide users with total control of the model. Users will be able to fine-tune settings such as temperature, top-p, repetition penalty, and max tokens, allowing for a more personalized experience. Additionally, the project will introduce a range of predefined model personalities as well as custom context, enabling users to choose from different response styles, tones, and behaviours that best suit their preferences and use cases. I will deliver an open-source web application with comprehensive documentation accompanying it, which includes instructions, usage guides, and explanations of how different parameters influence the model's output. I believe that this project will be a valuable addition to the Gemma ecosystem on Hugging Face.","date_archived":"2025-05-06T18:19:33.126411Z","id":"xYQnvnz5","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"},{"title":"Develop a Gemini Workspace in Postman","project_code_url":"https://www.postman.com/lorenzo-2698411/workspace/gemini-api","date_created":"2025-05-06T18:19:33.700284Z","tech_tags":["javascript","git","GitHub Actions","REST APIs","POSTMan","Gemini","Mock Servers","Postman Workspace"],"topic_tags":["testing","automation","documentation","CI/CD","Technical Writing","API Integration","LLMs","Cloud Development"],"status":"passed","program_slug":"2025","contributor_display_name":"Lorenzo Drudi","mentor_names":["Google DeepMind"],"abstract_short":"The goal of this project is to create a developer-friendly Postman Workspace for interacting with the Gemini API. This workspace will serve as a...","abstract_html":"The goal of this project is to create a developer-friendly Postman Workspace for interacting with the Gemini API. This workspace will serve as a central hub for exploring, integrating, and troubleshooting the Gemini API, providing developers with pre-built collections, test scripts, and documentation to streamline their workflow and reduce the learning curve. The workspace will feature well-documented API requests for key Gemini functionalities like text generation, chat, image generation, and code generation. It will also include pre-configured environments for testing and production, secure management of API keys, and a GitHub Action for automated updates.\n\nThe key deliverables for this project include:\n1. A comprehensive Postman Workspace with collections for various Gemini API features.\n2. Pre-configured environments for testing and production.\n3. Integrated, self-contained documentation and tutorials.\n4. Test scripts to validate API responses and handle edge cases.\n5. Mock servers to simulate API responses during local development.\n6. A GitHub Action to automate workspace updates based on API changes.\n\nThis solution will simplify the process for developers, improve productivity, and ensure the Postman Workspace remains up-to-date and reliable.","date_archived":"2025-05-06T18:19:33.700284Z","id":"naa1qGoe","organization_name":"Google DeepMind","organization_slug":"google-deepmind-sq"}]}