Monday, October 12, 2026
AI 인프라 · 뉴스 & 분석
홈 › 반도체·하드웨어 › 리포트
반도체·하드웨어 · 리포트

3D 칩을 높이 쌓는다고 해서 성능이나 수익이 반드시 향상되는 것은 아니며, 3D 집적의 상업적 확대 가능 여부는 본딩 정밀도뿐 아니라 수율과 양품 칩 하나당 생산 비용에 달려 있다.

고대역폭 메모리 및 첨단 패키징 구매자에게는 적층 높이가 아니라 양품 다이당 수율과 원가가 어떤 3D 설계가 양산에 이를지를 결정할 것이다.
업계 전문지Slicast · 2026년 10월 11일 04:27 UTC · 중국 · 출처: 钛媒体
중요도 40

*This article by Ban Dai Jun (半呆君) was licensed from TMTPost for publication.*

What if a chip isn't powerful enough? Could you stack another one on top?

Planar manufacturing is approaching its physical limits, and stacking chips vertically instead of spreading them out flat has become the shared strategy of leading manufacturers. TSMC is advancing chip-stacking technology. Intel's new server chips have already shipped. AMD has spent four years shipping products that stack cache. Huawei has used logic folding in its flagship smartphone.

But there is a problem that is easy to overlook: stacking chips does not automatically stack up profits. Each additional layer adds another manufacturing and connection checkpoint. After stacking, yield can fall rather than rise, and if a connection fails, the cost invested up to that point is scrapped with it.

The problem is also expensive to solve. Some processes handle an entire wafer at once. They are fast, but bad dies on the top and bottom layers drag each other down. Other processes first select known-good dies and then connect them one at a time. Each step costs more, but it can reduce losses from bad dies. This raises a counterintuitive question: why might a more expensive process ultimately produce cheaper, usable chips?

To answer it, we need to work through three questions in order: Why does stacking increase manufacturing risk? When does a more expensive process become the more economical choice? And what evidence shows that a new technology has moved from engineering breakthrough to commercial return?

**I. The taller the stack, the harder it is to make money**

Consider a simple case. Suppose you stack two chips to make a more powerful one. Each chip has a 95% yield. The chips' good or bad status is independent of each other, so the probability that both are good after stacking is not 95%. It is 95% × 95% = 90.25%.

This is the first hurdle of 3D integration. Each layer may look fine on its own, but once stacked, yield is eroded layer by layer. Real production is more complicated. With each added layer, beyond manufacturing the layer itself, you must align it with the other layers, connect and test them, and ensure the stack survives years of use. Every additional step multiplies the total by another number less than one.

The density of connections between chips (known in the industry as pitch) is important, since denser connections mean better performance. But density alone does not indicate mass-production capability. Mike Kelly, Vice President at Amkor, which handles packaging operations, gave an industry benchmark this year: the smallest pitch currently commercially viable is about 6 micrometers. The 1-micrometer pitch shown on roadmaps is unlikely to become a common mass-production choice in the near term. TSMC's trajectory is consistent with this. According to industry reporting summarizing its technology disclosures, its bonding pitch went from about 9 micrometers in 2023 to about 6 micrometers in 2025. A 4.5-micrometer pitch is expected around 2029. That is a planning target, not a reality yet.

*Yield is multiplication, not addition.*

So the answer to the first question is that the risk of stacking does not lie in whether any single layer was made well. It lies in whether every layer and every process step succeeds together. That raises the next question: if this is so difficult, why are companies still willing to pay for it?

**II. A more expensive process can produce cheaper chips**

There are two main ways to stack chips, representing two production strategies.

The first is wafer-to-wafer (W2W) bonding: two entire wafers are aligned and bonded in one operation, which is efficient. The drawback is that if a die on the top wafer is bad while the corresponding die on the bottom wafer is good, that pair is still scrapped. Defects jointly erode the final output.

The second is die-to-wafer (D2W) bonding: dies are tested first, and only confirmed good dies are placed onto the base wafer. The advantage is that bad dies never enter the stack. The cost is slower die-by-die processing and higher demands on equipment precision.

Neither approach is absolutely more advanced. One prioritizes overall wafer efficiency, while the other prioritizes the quality of the chips entering the stacking step. Which is more economical depends on whether the losses avoided from bad dies can cover the additional process costs.

I worked through the numbers using a set of illustrative parameters. The conclusion first: under these assumptions, the die-by-die known-good-die route has a higher cost per operation, but a lower cost per good stacked chip. The W2W route costs about $107.97, and the D2W route costs about $98.98, roughly 8.3% lower relative to the W2W cost.

Here are the calculations, for verification. The model inputs follow standard industry assumptions. Wafer manufacturing costs are $20,000 and $15,000 for the two wafers. Each wafer yields 600 chips. The die yield for each layer is 85%. The bonding pass rate is 98%, and the pass rate for subsequent steps is 99%. The W2W route costs $2,000 per wafer-pair bond, and the D2W route costs $8 per placement-and-bond operation.

For the W2W route, starting from the position level, output is 0.85 × 0.85 × 0.98 × 0.99, or about 70.1%. If either the top or bottom die at a given position is bad, that position is lost. In the D2W route, bad dies are stopped at the screening gate and never enter stacking. The stacking stage therefore only needs to pass the bonding and subsequent gates: 0.85 × 0.98 × 0.99, or about 82.5%. Measured on the same basis, the D2W route is still more than ten percentage points higher. One metric that matters only to specialists: good dies that have already entered bonding retain 97% after bonding. This 97% has a different denominator than the 70.1% figure, so the two should not be compared directly.

Finally, there is the break-even point. With all other parameters unchanged, the two routes reach parity when the per-operation placement-and-bonding cost for the D2W route rises to about $16.72. This figure holds only under these assumptions and must be recalculated if the parameters change. It is worth monitoring because it turns "how much is it worth to select known-good dies first" into a measurable quantity. To be clear, all of these figures are illustrative assumptions and do not correspond to any manufacturer's actual costs.

*Cost comparison of the two bonding routes*

The "bonding pass rate" assumed in these calculations is not determined by the bonding machine alone on a real production line. That leads to the next question: which steps determine whether this chain can run stably?

**III. The real bottleneck is not the bonding machine**

On October 1, Applied Materials and Besi announced an expanded partnership. Their collaboration, previously focused on die-to-wafer hybrid bonding, now extends to logic and memory integration, photonics, and a broader range of bonding platforms and interconnect architectures. Since their joint development center in Singapore in 2020, the two companies have produced what Applied Materials describes as the industry's first integrated die-to-wafer bonding system. Hybrid bonding is the step common to both approaches described above: without solder joints, dies polished to near-perfect flatness are bonded directly together. This expansion is evidence of deeper process collaboration. It is not an exclusive arrangement, and it does not indicate that revenue has shifted.

Viewed as a process chain, hybrid bonding involves five stages in sequence:

- **Cleaning and surface treatment:** remove particles and contaminants. Even minor contamination can disrupt the bonding interface.

- **Planarization and surface control:** bring the surfaces to be bonded to the required flatness. If the surface is out of spec, precise alignment is useless.

- **Precision alignment and placement:** position the upper and lower chips correctly. Any misalignment directly causes connection failure.

- **Bonding and annealing:** form a reliable connection at the interface. This determines how many defects remain.

- **Inspection, metrology, and feedback:** detect defects early and return the data to upstream steps. This determines whether yield can be sustained and stable.

The bonding machine is only one link in the chain. Research into failure mechanisms also points to the process as a whole. A single nanometer-scale particle can separate a large area of the interface that should have bonded. No matter how precise a bonding machine is, it cannot compensate for variation in upstream surface preparation and planarization.

*Hybrid bonding process chain: the opportunity is not limited to one machine.*

Following this chain, the industry opportunity breaks down into three questions that should not be conflated. Which stage's technical requirements are rising fastest? Which stage is hardest to substitute? And which company's related business is already generating revenue, rather than just technical positioning?

According to public disclosures, Applied Materials covers process platforms. Besi focuses on precision placement and bonding. EVG and SUSS MicroTec work on wafer bonding. ASMPT focuses on assembly and thermocompression bonding. KLA and Onto Innovation work on inspection and metrology. Piotech has positioned itself in hybrid bonding and supporting inspection and metrology equipment. However, rising technical requirements do not guarantee profit, and difficulty of substitution does not guarantee pricing power. Positioning is not the same as having substantial orders. These three distinctions mean the statement "profits have already migrated" cannot be made.

**IV. Logic goes first, memory waits**

The five stages in Section III determine whether mass production can be stable. But whether to adopt the technology also depends on a second test, which is economic. The same process capability produces different results on different products, depending on a comparison between the additional costs and the additional benefits.

For logic chips, benefits more easily cover costs. Take AMD's 3D V-Cache. It has been in mass production since 2022, roughly four years. By stacking a cache layer beside the processor, it reduces the number of trips to distant main memory. In certain workloads, the resulting performance gains justify the more complex manufacturing. TSMC's SoIC (System on Integrated Chips) stacking process is seeing continued expansion in capacity and customers. Intel's Foveros Direct has also appeared in reports about server chips.

For memory chips, older processes still meet current needs. HBM (High Bandwidth Memory) already achieves high bandwidth through multi-layer stacking. According to industry media reports, the maximum package height for HBM4 has increased from about 720 micrometers to 775 micrometers. The extra headroom allows 12-layer and 16-layer designs to continue using micro-bumps, the conventional small solder-ball connections. A TrendForce report in July said that Samsung and SK hynix are on different timelines for adopting hybrid bonding, and that 16-layer HBM4E or later generations may be the earliest window for it. I have not verified the exact wording and effective dates of the standards against primary documents, so they are presented here only as reported.

Connecting the memory case back to the break-even point calculated earlier: when older processes still meet height, yield, and cost requirements, a new process must demonstrate that its additional bonding cost can be offset by enough gains in layer count, bandwidth, or power efficiency.

Heat imposes a similar constraint. A simulation study of a 7-nanometer processor (published at IEEE ECTC 2020) found that, without thermal constraints, stacking raised the maximum junction temperature by about 12°C compared with a flat layout. When logic and memory were arranged in separate layers, the increase narrowed to about 6°C. These results apply to a specific structure and simulation conditions, but they help explain how the two product lines are divided. Cache layers generate little heat, so placing them beneath other layers does not consume much of the thermal budget. That is why "logic stacked on cache" reached mass production first. Stacking high-heat logic chips directly on another high-heat logic chip means heat compounds on heat, and so far this has appeared only in liquid-cooled data centers. Huawei's reported thermal figure of 300 W/cm² is a value disclosed in reporting. The test structure, boundary conditions, and scalability have yet to be verified.

As standards loosen, older processes can extend their life. Without cost parity, adoption is unlikely to happen.

*Adoption timelines for logic and memory: two tracks*

**V. Assessing a new technology: how far has the evidence progressed?**

The preceding analysis ultimately reduces to a chain of verification: the underlying principle is sound, the product structure is confirmed, performance gains are validated, manufacturing yield and cost meet requirements, stable volume delivery is achieved, and sustainable commercial returns follow. This is not a sequence every technology must follow in order. It is a yardstick for measuring how many steps remain before a technology makes money.

Here is how the yardstick applies to three recent headlines.

*Huawei's logic folding:* Structural evidence exists, as third-party teardowns have identified two stacked layers. Some third-party testing of performance exists, but the reductions reported in the paper (41% for computing, 58% for graphics, and 66% for AI) are based on the framing of Huawei's July update paper. Test conditions must be specified item by item, and the results have not been reproduced under the same conditions. There is no public data on yield or unit cost. Conclusion: the structural layer is established, the performance layer has partial third-party testing but no reproduction under matched conditions, and the economic layer remains blank. By this chain of verification, the technology is stuck at the transition from the second to the third layer. It has neither "proven only one layer" nor "been fully verified."

*Reports that NVIDIA will adopt Intel's Foveros:* This is still short of the structural layer. An analyst report relayed by several media outlets on October 9 is an opinion, not an order. There are no contracts and no product models. Formal customer relationships, packaging versions, order volumes, and production timelines all remain to be verified.

*Qualcomm and Huawei's patent agreement:* On October 6, Qualcomm issued a supplementary statement clarifying that the claim linking this agreement to logic folding is inaccurate. A single licensing agreement does not reach the level of "technology authorization."

Patents, principles, structural implementation, performance data, yield and cost, and commercial mass production are different tiers of evidence, and one cannot substitute for another.

This yardstick is worth taking with you. The next time you see CPO (co-packaged optics), silicon photonics, or any "next-generation packaging" claim, ask first: how far has the evidence progressed, and what will fill the missing layer?

So the endgame of 3D integration is not about who can reduce pitch the most. It is about who can integrate performance, energy efficiency, yield, cost, and reliable delivery into a repeatable engineering system.

You do not need to guess. Four observable indicators are enough to track:

1. Next steps in memory standards and manufacturer announcements. Distinguish between the standard text itself, media reporting, and manufacturer roadmaps.

2. Besi's Q3 report on October 22. Watch orders, order backlog, and management's comments on hybrid bonding demand. Customer counts should be taken from formal disclosures. The question raised in Section III, "who will first make money in cleaning and inspection," may be answered in financial reports and announcements like these. Current public evidence is not yet sufficient to draw a conclusion.

3. Yield, manufacturing cost, or batch stability for Huawei's logic folding. Public disclosure of any of these would be a key addition.

4. Third-party reproduction under matched conditions. Tests that specify the workload, frequency, voltage, ambient temperature, power measurement methodology, and sustained-load duration are far more persuasive than any secondhand account.

*The debate over 3D integration will ultimately be settled in the ledger of yield and cost.*

*(The author works in the industry supply chain. This article is for reference only and does not constitute investment advice. The cost model in this article uses illustrative assumptions and does not represent the actual costs, yields, or profit forecasts of any company.)*

*For more content, follow the TMTPost WeChat account (ID: taimeiti), or download the TMTPost app.*

원문 보기
3D 칩을 높이 쌓는다고 해서 성능이나 수익이 반드시 향상되는 것은 아니며, 3D… · Slicast