AI PCs Get Pricier: DDR4 Returns as Local LLMs Improve
Before AI PCs could fully establish themselves as the next major computing category, the AI boom began creating a problem for the hardware industry itself: the memory and storage components needed to build affordable computers are becoming substantially more expensive.
On October 9, 2026, GIGABYTE announced BIOS support for upcoming Intel LGA1700 processors across its entire B760 and H610 motherboard lineup, including both DDR4 and DDR5 models. The processors are expected to arrive in early 2027, offering existing platform owners another potential upgrade path without replacing their motherboards.
The decision reflects an uncomfortable market reality. AI infrastructure is consuming enormous quantities of memory, driving manufacturers to prioritize higher-margin data-center products while consumer PC builders face rising component costs. At the same time, global PC shipments have fallen sharply, suggesting that higher prices and inventory adjustments are putting pressure on demand.
Yet another development is moving in the opposite direction. Open-source software is making it possible to run increasingly large language models on ordinary gaming PCs by distributing workloads across GPU memory, system RAM, and CPU resources.
The result is a striking contradiction: AI is making computers more expensive to manufacture, while software advances are making the intelligence those computers can deliver increasingly accessible.
📉 The Memory Crisis Behind the PC Market Decline #
Rising memory prices are no longer confined to data-center procurement. They are affecting the component costs, retail prices, and upgrade decisions of ordinary PC users.
Why AI Demand Is Pushing Up Memory Costs #
AI servers require large quantities of high-bandwidth memory (HBM), server DRAM, and other high-performance components. As demand for AI infrastructure expands, memory manufacturers have incentives to allocate more production capacity to products with stronger margins.
This creates pressure across the broader memory market. Even when consumer DRAM and SSDs are not directly used in AI accelerators, their supply and pricing can be affected by manufacturing capacity allocation and competing demand.
Omdia reported on October 9 that the cost share of DRAM and SSDs in PCs had increased from a typical level of approximately 15% to nearly 40%. The research firm attributed the shift to substantial price increases across memory and storage components.
For more context on the infrastructure side of this trend, see KAD’s analysis of AI server price increases driven by HBM and DRAM shortages. The same pressure affecting data-center economics is increasingly relevant to consumer hardware.
When memory and storage consume a much larger share of a PC’s bill of materials (BOM), manufacturers have fewer options for maintaining the same specifications at the same retail price.
They can raise prices, absorb some of the additional cost, reduce specifications, or extend the lifespan of existing hardware platforms. The renewed interest in DDR4 compatibility is an example of the last two strategies.
Q3 2026 PC Shipments Fall Sharply #
The third quarter is traditionally an important period for PC manufacturers because of back-to-school purchases, commercial refresh cycles, and preparations for holiday demand. In 2026, however, the market experienced a sharp year-over-year contraction.
Two market research firms published preliminary estimates on October 9:
| Market tracker | Q3 2026 shipments | Year-over-year change |
|---|---|---|
| IDC | 62.7 million units | -20.1% |
| Omdia | 58.1 million units | -21.2% |
The totals differ because the firms use their respective market definitions and tracking methodologies. Both nevertheless indicate a severe decline in worldwide PC shipments.
Sources: IDC preliminary PC shipment results and Omdia’s Q3 2026 market analysis.
The downturn reflects several overlapping factors.
A difficult year-over-year comparison. Demand associated with the end of Windows 10 support helped pull purchases forward into 2025, raising the comparative baseline for 2026.
Early inventory accumulation. Vendors and distribution channels increased purchases earlier in 2026 to get ahead of anticipated component-price increases. That activity pulled some future demand into earlier quarters.
Higher retail prices. As increased procurement costs reached finished products, some consumers and businesses postponed purchases rather than paying substantially more for comparable hardware.
Weaker seasonal demand. Retailers entered the third quarter with inventory to clear, limiting the usual back-to-school boost.
Apple was affected as well. IDC estimated that Mac shipments fell approximately 11.3% year over year in Q3 2026, although Apple’s market share increased because its decline was smaller than that of the overall market.
These figures measure shipments rather than completed retail sales, and they do not establish that memory prices alone caused the contraction. The broader evidence points to a combination of component shortages, elevated prices, inventory adjustments, and weaker demand.
Why More Expensive PCs Become Harder to Sell #
PC manufacturers have already raised prices, but higher component costs can continue to compress margins.
When a conventional laptop becomes significantly more expensive without delivering a proportional increase in practical performance, buyers have a stronger incentive to delay upgrades. Businesses may extend hardware replacement cycles, while enthusiasts may reuse existing motherboards, RAM, storage, and cooling systems.
This creates a difficult feedback loop: higher component costs raise finished-product prices, weaker demand increases inventory pressure, and manufacturers become more reluctant to commit to new production.
The market response is increasingly shifting from maximizing specifications to preserving value. Older but functional platforms can become commercially attractive again when the cost of replacing an entire system rises sharply.
🔧 DDR4 Returns: Why Gigabyte Is Extending LGA1700 Support #
DDR4 was introduced in 2014 and has since been succeeded by DDR5 in newer mainstream platforms. Under normal market conditions, it would be natural for manufacturers to concentrate on newer memory standards.
But the current price gap has changed the calculation.
Gigabyte’s B760 and H610 BIOS Update #
In its October 9 announcement, GIGABYTE confirmed BIOS support for upcoming Intel processors using the LGA1700 socket across its entire B760 and H610 motherboard lineup, covering both DDR4 and DDR5 variants.
The new processors are expected to launch in early 2027. Their final model names, architecture, specifications, and performance characteristics have not yet been fully disclosed in the announcement.
For users who already own a compatible GIGABYTE motherboard, the intended upgrade path is straightforward: update the BIOS when the appropriate version is available, install the supported processor, and retain the existing motherboard and memory.
See GIGABYTE’s official LGA1700 compatibility announcement for the manufacturer’s stated support plan.
This approach can reduce the cost of a platform upgrade because users do not necessarily need to replace several major components simultaneously.
DDR4 vs. DDR5: The Price Gap Matters #
DDR5 remains the more advanced memory standard in terms of bandwidth and power efficiency. However, a substantial price difference can outweigh those benefits for buyers whose workloads do not require the newer standard’s performance.
In an October 9 report, The Verge documented a US retail price comparison between a 32GB Corsair Vengeance DDR5 kit at $620 and a corresponding DDR4 kit at $260.
That is a $360 difference for the memory alone. Retail prices fluctuate and individual listings do not establish market-wide averages, but the gap illustrates how strongly component costs can influence buying decisions.
For existing DDR4 users, reusing their current memory may be more economical than purchasing a new motherboard, new RAM, and a new processor together.
There are important limitations:
- DDR4 and DDR5 require different motherboard implementations and are not interchangeable in the same memory slot.
- An older motherboard needs a BIOS version that explicitly supports the new processor.
- The upcoming Intel processor specifications have not yet established how much performance users should expect.
- Compatibility does not necessarily imply that every existing system will receive identical performance or feature support.
The value of the update is therefore primarily about preserving an affordable upgrade path, not demonstrating that DDR4 is technically superior to DDR5.
AMD Is Also Revisiting an Older Platform #
Intel is not the only company benefiting from the continued relevance of older hardware.
GIGABYTE previously announced additional AMD AM4 motherboard options paired with the reintroduction of the Ryzen 7 5800X3D.
AM4 remains attractive to users who already own compatible DDR4 memory and want a processor upgrade without rebuilding their entire systems.
These developments suggest that platform longevity can become a competitive advantage when component prices rise. A mature motherboard and memory ecosystem may be more appealing than moving to a newer architecture whose performance benefits do not justify the additional cost for every user.
The broader lesson is that technology generations do not disappear simply because a successor exists. Market conditions determine when replacing older hardware makes economic sense.
🤖 Strata: Running a 125B-Parameter Model on a Gaming PC #
While manufacturers seek ways to make hardware more affordable, open-source developers are improving what existing hardware can accomplish.
A particularly interesting example is Strata, an open-source inference project designed to run the 125-billion-parameter Qwen3.8-Flash-Next model on consumer gaming PCs.
The project’s GitHub repository publishes benchmark results for systems with relatively modest graphics memory, including an NVIDIA RTX 5070 with 12GB of VRAM, a Ryzen 5 7600 processor, and 64GB of system RAM.
The central idea is simple but technically important: the entire model does not have to reside in GPU VRAM for a system to generate useful output.
Why a 125B Model Can Run on Limited VRAM #
A 125-billion-parameter model would require far more than 12GB of memory if every parameter were represented at conventional full precision.
Strata takes advantage of several techniques to make inference practical on consumer hardware.
Mixture-of-Experts architecture. The model uses an MoE design in which only a subset of expert parameters is activated for a given token. Its total parameter count therefore does not directly translate into the amount of computation needed for every generation step.
Quantization. Lower-bit representations reduce the amount of memory required to store model weights. More aggressive quantization allows larger models to fit within the available combination of VRAM and system RAM, although it can affect output quality.
Heterogeneous execution. Strata distributes work across the GPU, CPU, and system memory. Frequently used expert weights can be kept in VRAM, while additional model data resides in system RAM and CPU resources contribute to execution when appropriate.
Speculative decoding. Smaller auxiliary components predict candidate tokens that the main model can verify in batches. When predictions are accepted, the system can reduce the number of sequential generation steps.
This is a form of software-hardware co-design: the inference engine adapts its execution strategy to the memory and compute resources available in the PC rather than assuming that the whole model fits on a single GPU.
For another example of this approach, KAD’s article on FreeToken running large MoE models with consumer GPUs and system memory discusses how expert caching and heterogeneous execution can expand the practical size of locally runnable models.
RTX 5070 Benchmark Results #
According to Strata’s published results for an RTX 5070 with 12GB of VRAM and 64GB of system RAM, generation speed varies substantially with quantization settings.
| Quantization configuration | Reported generation speed |
|---|---|
Q2_0 |
Up to 94 tokens/s |
IQ2_XS |
Approximately 79 tokens/s |
IQ3_XXS |
Approximately 62 tokens/s |
IQ3_S |
Approximately 53 tokens/s |
| Coder configuration | Approximately 55 tokens/s |
These numbers are reported by the project maintainers, not independently verified universal performance figures. Results depend on the particular engine version, prompt length, generated output length, background workload, and configuration.
The repository distinguishes prompt-processing speed from token-generation speed. Its results use different workloads and inference settings, so generation speed should not be confused with how quickly the model can ingest a long document.
Even with those caveats, the results demonstrate how far inference software can stretch hardware that would previously have been considered too constrained for a model of this size.
At 94 tokens per second, output appears much faster than most people can read comfortably. The more important engineering achievement is that a model with 125 billion total parameters can operate on a consumer PC whose discrete GPU has only 12GB of VRAM.
Quantization Is a Trade-Off, Not a Free Upgrade #
Lower-bit quantization reduces memory requirements, but it introduces compromises.
A highly compressed representation can alter model outputs, reduce reliability on certain reasoning tasks, and affect coding or language-generation quality. The severity depends on the model, quantization method, calibration data, and inference implementation.
The speed results therefore need to be evaluated alongside output quality, task completion rates, and the demands of the intended application.
A more compressed model is not automatically the better model. It may provide enough quality for summarization or routine coding assistance, yet struggle with demanding reasoning tasks compared with a less aggressively quantized configuration.
Strata’s performance results should be viewed as one point in a broader trade-off among memory footprint, speed, and fidelity.
Speculative Decoding and Memory Hierarchy Optimization #
The project also uses speculative decoding to reduce serial generation work.
A smaller auxiliary model or prediction component proposes candidate tokens. The main model evaluates these candidates together, accepting compatible predictions and correcting rejected ones.
Strata’s documentation reports acceleration of roughly 1.6× to 1.8× in supported configurations, although the actual gain depends on the workload and prediction acceptance rate.
Speculative decoding complements quantization and expert caching, but does not replace either technique. It reduces the number of sequential decoding steps; it does not restore precision lost through quantization or eliminate the memory bandwidth cost of accessing model weights.
The broader engineering lesson is that performance improvements often come from combining several moderate optimizations rather than relying on one dramatic hardware upgrade.
🖥️ System RAM Is Just as Important as VRAM #
The Strata example also highlights why GPU memory capacity is only one part of the local inference equation.
An RTX 5070 with 12GB of VRAM may be sufficient to execute some parts of a large model, but system RAM provides the capacity needed to keep the remaining data available.
Why 64GB of RAM Changes the Experience #
With 64GB of system RAM, the tested system can accommodate a large amount of model data alongside the operating system, applications, and inference runtime.
By contrast, a 32GB configuration has less headroom. Depending on the model variant and the operating environment, users may need to select a more aggressively quantized or reduced model configuration.
Memory usage also depends on context length, KV cache requirements, runtime buffers, and whether additional applications are running.
This explains why a PC with the same GPU can deliver a substantially different local AI experience depending on its system memory configuration.
Startup Time and Multitasking Costs #
Loading tens of gigabytes of model data is not instantaneous.
Strata’s documentation notes that its initial loading process can cause significant system activity and temporary slowdown, with startup waits of approximately one to three minutes on some configurations.
Long inputs also require prompt processing before generation can begin. A high token-generation speed does not mean a long document is processed instantaneously.
Other costs include:
- Memory consumed by multiple models or applications.
- CPU contention when some inference operations execute outside VRAM.
- Storage access during initialization or data loading.
- Potential reductions in responsiveness while large allocations are being established.
- Quantization-related changes in model output quality.
The practical value of local AI therefore depends on more than a single throughput number. Startup time, sustained performance, background multitasking, prompt-processing speed, and output quality all determine whether an implementation is useful for daily work.
What Existing PCs Can Achieve #
Strata demonstrates that a gaming PC does not need to be replaced merely because its GPU cannot hold an entire large model.
By using system RAM as an additional capacity tier and assigning work to different compute resources, an existing configuration may become capable of workloads that previously appeared out of reach.
This does not mean every older PC can run a 125B model effectively. A compatible GPU, sufficient system RAM, a supported operating system, and a suitable inference backend are still required.
However, it reveals a practical alternative to hardware replacement: improve the efficiency with which existing hardware is used.
💰 AI Is Making Intelligence Cheaper While Hardware Gets More Expensive #
The PC market is facing two opposing economic trends.
On one side, the cost of accessing capable language models has fallen rapidly as inference software improves, competition increases, and new model architectures reduce the resources required for a given task.
On the other, the physical infrastructure supporting those models—including DRAM, HBM, GPUs, and SSDs—is experiencing supply constraints and higher prices.
A chart attributed to Goldman Sachs and circulated in investor commentary has illustrated the speed of AI model-cost reductions relative to historical PC price declines, comparing roughly three years of falling LLM costs with a much longer period of PC price reductions. Such comparisons depend on the chosen price index and quality adjustments, but they capture the direction of change.
The cost of consuming intelligence and the cost of manufacturing the hardware that serves it do not have to move in the same direction.
Software Optimization Extends Hardware Lifespans #
When inference software uses quantization, speculative decoding, intelligent caching, and heterogeneous execution effectively, a PC can deliver more useful AI performance without a corresponding increase in silicon or memory capacity.
This has several implications:
- Consumers can retain older hardware for longer while still accessing capable local models.
- Developers can prototype and test agents locally before investing in larger workstations or servers.
- Existing VRAM and system memory can be used more efficiently through better scheduling and caching.
- More workloads can run locally, potentially reducing dependence on recurring cloud inference costs.
These benefits are not unlimited. Software cannot erase the bandwidth and capacity constraints of physical hardware. But it can make the difference between a configuration that cannot run a workload and one that runs it with acceptable performance.
Choosing Between Hardware Upgrades and Software Optimization #
For users deciding whether to buy a new PC, upgrade memory, or optimize local inference, the right approach depends on the actual bottleneck.
| Bottleneck | Potential response |
|---|---|
| Insufficient system RAM | Add compatible memory if the platform supports expansion |
| Limited GPU VRAM | Use quantization, expert caching, model offloading, or a GPU with more VRAM |
| Slow token generation | Evaluate the inference backend, quantization format, GPU utilization, and speculative decoding |
| Excessive prompt-processing time | Optimize context length, prompt processing, and model configuration |
| High hardware acquisition cost | Reuse compatible components, consider mature platforms, or use cloud inference selectively |
| Poor output quality | Use a less aggressive quantization configuration or a more capable model |
The key is to identify the limiting resource before spending money. A system may need more RAM, but it may instead be constrained by GPU compute, memory bandwidth, a poorly optimized backend, or excessive context length.
The cheapest solution is not always the lowest-cost component. It is the change that removes the actual bottleneck with the least overall expense.
✅ Conclusion: The Best PC Upgrade May Be Better Software #
The AI boom is creating an unusual situation for personal computing: the demand for AI infrastructure is raising memory and component costs just as software advances are making powerful AI workloads more accessible on ordinary PCs.
GIGABYTE’s renewed support for upcoming LGA1700 processors across its DDR4 and DDR5 B760 and H610 motherboard families reflects the growing value of platform longevity. For existing users, preserving a compatible motherboard and reusing memory can reduce the cost of an upgrade.
Meanwhile, Strata demonstrates how quantization, MoE execution, CPU-GPU cooperation, and speculative decoding can make a 125-billion-parameter model usable on a gaming PC with 12GB of VRAM and 64GB of system RAM.
Three conclusions stand out:
- Memory prices influence the entire PC market: AI-driven demand is changing component economics and encouraging users to extend the lives of existing systems.
- Mature hardware remains useful: DDR4 platforms can offer an economical upgrade path when compatibility and workload requirements align.
- Software determines how much hardware can achieve: Better inference engines can unlock workloads that once seemed impractical without a hardware replacement.
AI may continue making capable intelligence cheaper to access, but building and upgrading the devices that run it will remain sensitive to semiconductor supply and memory costs.
For PC users, the most effective strategy is increasingly twofold: preserve useful hardware where possible and optimize software around its real constraints. In an environment where a single memory kit can dramatically affect a build budget, getting more from an existing computer may be just as valuable as buying a faster one.