Summary
Hello, readers interested in cloud and AI infrastructure! Today, Google Cloud (Google Cloud) 7th generation TPU ‘Ironwood TPU (Ironwood) General Availability and The new ARM-based VM group Axion CPU (Axion) instance Let’s kindly unravel the news of the preview release in the form of a blog.
We will guide you by reflecting all the conditions for writing a blog (Markdown, table and list, including images and links).
This announcement is not just a data center upgrade. The transition from the age of ‘training’ to ‘inference’ of the AI model to ‘inference’It can be seen as a turning point clearly showing.
It is particularly noteworthy in that it is an infrastructure innovation designed to enable more efficient business customers to deal with increasingly inferior workloads (such as chatbot responses, agent-based applications, etc.).

Key presentations
Ironwood TPU
Source: Google Cloud Blog (CC0 – Similar License) – Ironwood board photo
- Google Cloud is Ironwood 7th generation TPU, compared to the existing TPU V5P and TPU V6E (Trillium) 10 times, said that it provides 4 times more performance improvement. Google Cloud+2insidehpc.com+2
- Not only training, but also for low-latency and high-throughput AI inference workloads, he emphasized that it was designed in line with the **inference age**.. blog.google+1
- SuperPod configuration is possible with up to **9,216 chips (TPU chips)**, and a shared high-bandwidth memory (HBM) 1.77 PB is supported and super-fast inter-chip interconnect (ICI). It is characterized by implementing networking 9.6TB/s.. Google Cloud+1
- By applying optical circuit switching (OCS) technology, real-time path reconfiguration is possible even in case of network failure, and large workload operation is possible without interruption of service.. The register +1
Axion VM Instances
Source: Google Cloud Blog – Axion VM example image
- In addition to AI-specialized accelerators (Ironwood), Daily Universal Computing WorkloadsARM Neoverse-based customized CPU group ‘AXION’ has been released to efficiently handle data preparation, collection, web service backend, etc.. Google Cloud+1
- Among them N4A (Preview) Instances are the largest compared to their class x86 VM 2x better price/performance ratioThey are said to provide. Google Cloud+1
- Also, **C4A Metal (Bare Metal Instance, Preview will be announced soon)** has been announced, designed to respond to Android development, in-vehicle systems, software with licensing constraints, complex simulations, etc.. Google Cloud+1
- In this way, by combining Ironwood and Axion’s instances, from learning to inference and operation Full-Stack AI HyperComputer Google Cloud’s strategy is to provide a form of integrated infrastructure.. Google Cloud+1
Why is the ‘AI Era of Inference’ now?
- The recent AI model is not just learning, but Real-time service for millions, tens of millions of users It is rapidly shifting to the focus of inferences provided in the form (e.g. chatbots, agents, customized creation AI).
- So, as much as ‘learning a model’, ‘inference to serve the learned model quickly and efficiently’ has become a key requirement for the company.
- Accordingly, the infrastructure blueprint Latency, Throughput, Cost-Efficiency, and Operating Efficiency (Ops Efficiency) It is being reorganized in the side.
- In line with these changes, Google Cloud is ‘co-designed from hardware to software’ through Ironwood TPU and Axion CPU instances. Infrastructure optimized for inference workloadsIt is said that we will provide. Google Cloud+1
Summary of main features and specifications
The table below summarizes the published specifications and features.
| Items | Features and descriptions |
|---|---|
| Generation | 7th generation TPU, Ironwood (TPU V7) |
| Preparing for improvement | Up to 10 times compared to TPU V5P, more than 4 times per chip compared to TPU V6E Google Cloud+1 |
| Scale | Superfords with up to 9,216 chips → expandable to hundreds of thousands The register +1 |
| Memory | Shared high bandwidth memory (HBM) 1.77 PB scale support Google Cloud+1 |
| networking | Inter-Chip Interconnect (ICI) 9.6 TB/s, Optical Circuit Switching (OCS) technology applied The Register |
| universal infrastructure | Axion Instances (N4A, C4A, C4A Metal) are cost-effective for universal workloads Google Cloud |
| Software Integration | VLLM, MaxText framework support, integrated Google Kubernetes Engine (GKE) cluster director function Google Cloud+1 |
Real Customers and Use Cases
- Anthropic, an AI startup, is based on Ironwood’s price-performance ratio. Plan to use up to 1 million TPUshas been published. anthropic+1
- In addition, the image processing platform Vimeo and data intelligence company ZooMinfo have also tested and operated AXION instances and were evaluated to have significantly improved cost-effectiveness.. Google Cloud
- These customer examples are not merely technical presentations, but Reduce your cloud infrastructure costs and Improve operational efficiencyIt shows that it is leading to
Benefits that a company can benefit from this announcement
- Response to low latency and high throughput inference workloads: Reduce performance bottlenecks in real-time user responses, agent-based applications, and more.
- Improve cost efficiency: Combining high-performance TPU and general-purpose instances, you can design optimized cost structures based on your workloads.
- Improve operation and scalability: In the form of a superford, it can be expanded from thousands to tens of thousands of chips, and the operational burden is reduced with management functions integrated with GKE, etc.
- Securing future responsiveness: It can be used as an infrastructure in response to a surge in AI model structure (e.g., mixed expert MOE, agent workflow, etc.) and AI usage.
Things to consider and tips
- Because the announcement is ‘officially released within a few weeks (with GA)’, the actual service availability and region may be limited.. Google Cloud
- It is important to plan which instances to use first, taking into account the characteristics of workloads (e.g., learning-oriented vs reasoning) and the current infrastructure (e.g. x86-based VMs, GPU-based accelerators, etc.).
- To minimize cost/operation risk Pilot ProjectI Trial useWe recommend that you review first.
- If you analyze the software stack support (JAX, PyTorch, MaxText, etc.) and the difficulty of the migration in advance, the introduction process will be smoother.
Conclusion
This announcement does not mean that Google Cloud is simply ‘faster hardware’ A paradigm turning point in AI infrastructureThis is a meaningful event.
In other words, it is also a sign that the shift to ‘training-centered → inference-focused’ and ‘integrated infrastructure that combines general CPU-based workloads and AI accelerators’ is in full swing.
From the company’s point of view, it is necessary to keep an eye on this technology trend and think about how to align it with its AI applications and cloud operation strategies in the future.
Now, based on the published content, I would like you to think together about what the company’s meaning will have and what strategy you will be able to come up with.
Next, we have prepared the most frequently asked questions (FAQs), so take a look!
FAQ
Q1. Should Ironwood TPU replace existing GPU-based cloud instances?
A1. No. Ironwood is a high-performance AI accelerator and is especially suitable for large-scale model learning and low-latency and high-throughput inference workloads. However, not all workloads are applicable. For example, experimental model training, small data processing, general CPU-based applications, etc. may be sufficient with existing GPUs or x86 instances. The best strategy is to first analyze the workload characteristics, Ironwood + Universal CPU (Axion, etc.) It’s just assessing if the combination offers an advantage.
Q2. What kind of companies are Axion VMs (N4A, C4A, etc.) particularly advantageous?
A2. Axion VM is Universal Computing WorkloadIt is an optimized instance. Examples include microservices, container-based applications, open source databases, data acquisition and preprocessing, web hosting, development/test environments, and more. Also, compared to the X86-based VM Up to 2x the price/performance ratioGoogle announced. Therefore, axion can be very advantageous for backend operations, data pipelines, and app servers, not AI model inference.
Q3. What preparations are needed when considering the introduction of Ironwood?
A3. Here are some things to consider when introducing:
- Workload analysis: Be sure to grasp the model size, inference frequency, latency, and throughput.
- Software compatibility: Check for compatibility with the relevant framework such as JAX, PyTorch, MaxText, etc.. Google Cloud+1
- Understanding the cost structure: Compare TPU and VM instances per hour, scalability, maintenance costs, and more.
- Establishing a Migration Plan: When switching from the existing GPU/CPU infrastructure to TPU, prepare data movement, code refactoring, and testing procedures.
- Pilot/Beta Test: It is recommended to test on a small scale first and verify the actual performance improvement and stability.
Q4. What does this technology presentation mean to the cloud market or to competitors?
A4. This presentation has the following meanings:
- In terms of cloud infrastructure Custom Silicone It is a sign that the strategy is becoming more important. Google is trying to differentiate in the GPU-oriented market by using its own AI accelerators through Ironwood.. Constellation Research Inc.+1
- For competitors (e.g. NVIDIA, Amazon Web Services, etc.), it means that the challenge of ‘infrastructure innovation to respond to AI inference workloads’ has grown even more.
- ‘It’s not enough to just use more GPU instances anymore, Optimizing inference, optimizing cost, and improving operational efficiencyThere is a high possibility that the perception that it should be the key to the selection of infrastructure will spread.
Q5. What is the meaning of a domestic company or in the Korean market?
A5. In Korean companies and markets, it has the following meanings:
- If a domestic company is preparing AI services (e.g., chatbot, generated AI, agent service), the latest TPU and VM options provided by Google Cloud, a global infrastructure provider, are Possibility of domestic regional useYou may be subject to review with this in mind.
- In particular, if ‘low delay (network latency)’, ‘mass user response’, and ‘high cost-effectiveness’ are required, the Ironwood-Axion combination can be competitive.
- However, you must also look at local factors to consider when introducing infrastructure, such as domestic regulation, data center location, legal/security requirements, and Korean language support.
- Finally, it can be used as an opportunity to reconsider ‘would it be better to build our own infrastructure to train a specific model, or to use the latest cloud infrastructure?’
Note link
- Google Cloud Blog: Announcing ironwood tpus general availability and new Axion VMS to power the age of inference – Google Cloud
- Reuters: “Google Launches New Ironwood Chip to Speed AI Applications”
- Techradar: ‘Anthropic Signs Multi-Billion Dollar Google Deal… Up to a Million Tpus’




댓글 남기기