<?xml version="1.0" encoding="UTF-8"?><rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title>Google Developers Blog</title><link>https://developers.googleblog.com</link><atom:link href="http://rss.144-124-237-35.sslip.io/google/developers/en" rel="self" type="application/rss+xml"></atom:link><description>Google Developers Blog - Powered by AtomRSS</description><generator>AtomRSS</generator><webMaster>contact@atomgroup.dev (AtomRSS)</webMaster><language>en</language><lastBuildDate>Sat, 08 Aug 2026 05:41:36 GMT</lastBuildDate><ttl>5</ttl><item><title>Agent Plugins package your skills, tools, and more</title><description>Agent Plugins 1.0.0 is a new, vendor-neutral directory specification—backed by Google, Amazon, Microsoft, and others—for packaging Agent Skills and MCP servers into a single portable unit. By standardizing the manifest (plugin.json) and utilizing a fixed directory layout, it eliminates the need for developers to maintain separate wrappers or configurations to support different AI coding agents and IDEs. Google has officially joined as a Core Maintainer and already rolled out support in the Agents CLI and Data Agent Kit, allowing developers to start building and distributing interoperable plugins today.</description><link>https://developers.googleblog.com/en/agent-plugins-package-your-skills-tools-and-more/</link><guid isPermaLink="false">https://developers.googleblog.com/en/agent-plugins-package-your-skills-tools-and-more/</guid><pubDate>Wed, 05 Aug 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item><item><title>Scaling AI Agent Infrastructure with the MCP Stateless updates</title><description>The 2026-07-28 Model Context Protocol (MCP) specification replaces legacy stateful constraints with a fully stateless core, enabling cloud-native horizontal scaling, serverless deployments, and standard round-robin load balancing. This architectural shift introduces standardized HTTP headers for efficient routing without deep packet inspection, caching controls, and Multi Round-Trip Requests (MRTR) to handle interactive and long-running tasks without blocking connections. Developers can immediately begin migrating their agentic applications to this highly scalable infrastructure using the newly available beta SDKs for Python, TypeScript, Go, and C#.</description><link>https://developers.googleblog.com/en/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates/</link><guid isPermaLink="false">https://developers.googleblog.com/en/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates/</guid><pubDate>Tue, 04 Aug 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item><item><title>A unified API for AI model routing</title><description>Google Cloud API Gateway now offers a model routing feature in Public Preview, allowing developers to dynamically route traffic to models like Gemini, Claude, or OpenAI OSS-GPT without hardcoding endpoints or managing open-source proxies. Developers can easily configure these routing rules directly within their OpenAPI 3.x specifications by mapping virtual model names to specific backend targets on a shared host. Once deployed, the Gateway acts as a serverless ingress layer that accepts standard OpenAI-compatible requests, automatically transcodes the payload to the native schema of the target model, and routes the traffic on the fly.</description><link>https://developers.googleblog.com/en/a-unified-api-for-ai-model-routing/</link><guid isPermaLink="false">https://developers.googleblog.com/en/a-unified-api-for-ai-model-routing/</guid><pubDate>Mon, 03 Aug 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item><item><title>Scaling real-time AI agents with session-aware load balancing</title><description>Real-time AI agents break traditional request-response load balancing paradigms because they rely on long-lived, stateful bidirectional streams that obscure true server capacity. To solve this, developers must implement application-level session tracking directly within the runtime to accurately measure the committed concurrent workload of active conversations. By feeding these precise session counts alongside standard CPU utilization metrics into a hybrid routing algorithm, infrastructure can effectively distribute stateful AI traffic and prevent individual backend bottlenecks.</description><link>https://developers.googleblog.com/en/scaling-real-time-ai-agents-with-session-aware-load-balancing/</link><guid isPermaLink="false">https://developers.googleblog.com/en/scaling-real-time-ai-agents-with-session-aware-load-balancing/</guid><pubDate>Sun, 02 Aug 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item><item><title>Enable on-demand expertise with Agent Skills in Genkit Go</title><description>To prevent context window bloat and reduce token consumption, Genkit Go introduces Agent Skills based on a progressive disclosure architecture. Developers can package specialized instructions, scripts, and references into modular SKILL.md bundles where only the frontmatter metadata is initially exposed to the agent&#39;s system prompt. When a task matches the skill&#39;s description, Genkit&#39;s middleware dynamically loads the full instruction body and associated assets, ensuring the model accesses precise workflows exactly when needed.</description><link>https://developers.googleblog.com/en/enable-on-demand-expertise-with-agent-skills-in-genkit-go/</link><guid isPermaLink="false">https://developers.googleblog.com/en/enable-on-demand-expertise-with-agent-skills-in-genkit-go/</guid><pubDate>Thu, 30 Jul 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item><item><title>Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA</title><description>Agent Platform&#39;s evaluation service is now generally available, providing developers with a unified engine to measure agent quality consistently across local development experiments and live production traffic. You can evaluate agents using over 20 pre-built metrics, DeepMind-backed adaptive rubrics, or custom code-based and LLM-as-a-judge metrics stored in a centralized, versioned registry. The service integrates directly into existing workflows via the Agent Platform SDK, agents-cli, and ADK, offering built-in user and environment simulators to automate complex multi-turn testing and streamline CI pipelines.</description><link>https://developers.googleblog.com/en/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga/</link><guid isPermaLink="false">https://developers.googleblog.com/en/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga/</guid><pubDate>Thu, 30 Jul 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item><item><title>How to use Google microbenchmarks for evaluating TPU performance</title><description>Google&#39;s open-source TPU microbenchmark suite provides developers with granular performance metrics across Network, Compute, HBM, Host Transfer, and Attention components to validate real-world hardware capabilities. By leveraging these benchmarks to establish a Roofline model, engineers can accurately diagnose whether their machine learning workloads are compute-, memory-, or network-bound. This empirical baseline directly guides targeted software optimizations—such as kernel tuning, mesh sharding, and rematerialization—to maximize hardware utilization for large-scale model deployments.</description><link>https://developers.googleblog.com/en/how-to-use-google-microbenchmarks-for-evaluating-tpu-performance/</link><guid isPermaLink="false">https://developers.googleblog.com/en/how-to-use-google-microbenchmarks-for-evaluating-tpu-performance/</guid><pubDate>Wed, 29 Jul 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item><item><title>Run Ray on TPU, Part 2: Ray AI libraries</title><description>This second installment explores how Ray’s higher-level libraries—Serve, Data, and Train—abstract the complexities of running AI workloads on Google&#39;s TPU slices. Ray Serve uses a simple topology configuration to correctly gang-schedule large multi-host models, while Ray Data eliminates data-loading bottlenecks by feeding accelerators directly with native JAX batches. Finally, JaxTrainer streamlines distributed training across TPUs by automatically handling cross-slice coordination, checkpointing, and fault tolerance.</description><link>https://developers.googleblog.com/en/run-ray-on-tpu-part-2-ray-ai-libraries/</link><guid isPermaLink="false">https://developers.googleblog.com/en/run-ray-on-tpu-part-2-ray-ai-libraries/</guid><pubDate>Thu, 23 Jul 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item><item><title>Scaling Agentic RL: High-Throughput Agentic Training with Tunix</title><description>Tunix is Google’s new JAX-native post-training library designed to eliminate TPU idling bottlenecks when training multi-turn, tool-using LLM reasoning agents. It maximizes hardware throughput by combining highly concurrent, asynchronous rollouts with a decoupled producer-consumer pipeline, ensuring the trainer is constantly fed even while agents wait on network I/O or environment steps. Additionally, Tunix provides plug-and-play abstractions and continuous macro-level profiling, allowing developers to easily integrate custom open-source environments and optimize complex distributed workflows without massive code rewrites.</description><link>https://developers.googleblog.com/en/scaling-agentic-rl-high-throughput-agentic-training-with-tunix/</link><guid isPermaLink="false">https://developers.googleblog.com/en/scaling-agentic-rl-high-throughput-agentic-training-with-tunix/</guid><pubDate>Mon, 20 Jul 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item><item><title>Run Ray on TPU, Part 1: The foundations</title><description>Ray 2.55 introduces official, first-class support for Google Cloud TPUs, enabling developers to run distributed Python workloads on Google&#39;s accelerators using the familiar Ray task-and-actor APIs. To handle the strict networking requirement of keeping multi-host TPU &quot;slices&quot; together over their Inter-Chip Interconnect (ICI), the KubeRay Operator on GKE automatically provisions and labels the underlying hardware layout. Ray Core utilizes these labels via its slice_placement_group() primitive to atomically reserve complete slices, allowing developers to deploy jobs through KubeRay, Ray Train, or Ray Serve simply by declaring a hardware topology (like &quot;4x4&quot;) without writing custom placement code.</description><link>https://developers.googleblog.com/en/run-ray-on-tpu-part-1-the-foundations/</link><guid isPermaLink="false">https://developers.googleblog.com/en/run-ray-on-tpu-part-1-the-foundations/</guid><pubDate>Sun, 19 Jul 2026 16:00:00 GMT</pubDate><author>Google</author><category>AI</category></item></channel></rss>