Im Eui-cheol, SK hynix Solution AT vice president, delivers keynote

More memory, less computation: KV cache reuse gains traction

Agentic AI runs 24/7, driving demand for GPU and CPU server infrastructure

3D-stacked DRAM and HBF set to advance memory hierarchy

Im Eui-cheol, vice president and head of Solution AT at SK hynix, delivers a keynote address titled "In the AI Era, Memory Is What Matters Most" at the Herald Business Forum 2026, held Tuesday at the Dynasty Hall of the Shilla Hotel in Jung-gu, Seoul.
Im Eui-cheol, vice president and head of Solution AT at SK hynix, delivers a keynote address titled "In the AI Era, Memory Is What Matters Most" at the Herald Business Forum 2026, held Tuesday at the Dynasty Hall of the Shilla Hotel in Jung-gu, Seoul.

"New algorithms to reduce memory usage are being developed every day. But that does not translate directly into a decline in overall memory demand."

Im Eui-cheol, vice president and head of Solution AT at SK hynix, delivered that message Tuesday at the Herald Business Forum 2026, held at the Dynasty Hall of the Shilla Hotel in Jung-gu, Seoul. Speaking under the theme "In the AI Era, Memory Is What Matters Most," Im argued that efficiency gains in AI will ultimately expand memory demand rather than shrink it — because even as memory consumption per token falls, lower costs drive total token usage higher.

"The reason we work so hard on this is that we want to use more tokens but do not have enough memory space to do so," Im said. "Even if we reduce the memory capacity per token, we will inevitably end up using more tokens overall." He added that the dynamic feeds on itself: "Because it gets cheaper, we use more — and as we store data to use it more efficiently, we need even more memory." That, he said, is the structure driving total demand upward.

Im said understanding the importance of memory in the AI era requires looking first at how large language models work. LLM inference is divided into two stages: prefill, which processes and understands the input, and decode, which generates tokens one by one.

"In the prefill stage, computation matters because once parameters are loaded, that data can be reused hundreds or thousands of times," Im said. "In the decode stage, the entire recursive LLM model must be read for every single word processed, which means overall LLM performance is determined by memory performance."

Speed matters, but so does capacity, he said. SK hynix has addressed the need for high bandwidth, large capacity and low latency through HBM, or high-bandwidth memory.

"We stacked memory dies to increase capacity within a small footprint, and drilled many I/O connections between dies using through-silicon vias to raise bandwidth," Im said. "At the same time, we minimized the distance between HBM and the compute unit to reduce power consumption."

A broader shift is underway in the AI industry toward using more memory to reduce the computational burden. "The concept of storing pre-computed results in memory and reusing them — rather than reprocessing them — is emerging," Im said. "The idea is to use more memory to improve the efficiency of the entire system." He added that the approach is fast becoming the industry standard.

The key piece of data enabling this is the KV cache, or Key-Value Cache, which stores key and value information calculated while processing earlier tokens so it can be reused in subsequent operations. As the volume of tokens processed and data stored grows, the importance of the entire memory hierarchy rises — not just HBM, but also server DRAM, SSDs and network storage.

Im said agentic AI will accelerate the expansion of memory demand even further. Unlike conventional chatbots, agentic AI can be given a task and will carry it out autonomously, breaking free from the pace at which a human reads a response and types the next question.

"When people use AI, they have to eat, sleep and take breaks," Im said. "With agentic AI, the user can hand off the work and go home." He added that agentic AI works seven days a week, 24 hours a day. "That means we need more data center infrastructure — not just GPUs, but CPUs as well," he said.

Im said AI services must achieve profitability to remain sustainable. While the value delivered by agentic AI has grown and service prices have risen accordingly, he said the priority is bringing down infrastructure costs.

SK hynix is developing next-generation memory technology to meet these varied demands. For the premium market where fast token processing is essential, the company is preparing 3D-stacked DRAM; for cost-sensitive services, it plans to supply high-bandwidth flash, or HBF.

On HBF, Im said the company is "thinking about providing bandwidth comparable to HBM and capacity comparable to NAND flash, to increase throughput at a lower price."

"Using memory effectively to improve the efficiency of computing systems has become critical," Im said in closing. "SK hynix is conducting research and development at the system level — looking not just at the device itself, but at how the entire architecture will evolve."


jeongwan@heraldcorp.com