llama.cpp Makes Adreno Prefill Up to 25% Faster Without New Silicon
llama.cpp release b10687 changes how two Qualcomm Adreno GPU generations route matrix multiplication. The project reports about 25% faster prefill for stock gpt-oss-20b on an Adreno X2-90 and about 9% for Gemma 3n E4B on an Adreno 740, without changing decode. The result is narrow, measured software efficiency—not a universal device speedup. Fresh filings show why that distinction matters financially. Marvell’s quarterly revenue rose 36.5% to $2.739 billion as data-center sales grew 46%, while Elastic’s subscription revenue rose 15% but subscription cost climbed 32%, lowering subscription gross margin from 82% to 80%. Federal Reserve data show commercial-bank loans and leases increased $35.3 billion in the week to August 19, but cannot identify AI financing. ZharfAI’s conclusion: the next useful unit of AI capacity may come from better dispatch on hardware already installed, yet value only becomes durable when workload gains survive across models and convert into healthier margins or cash flow.
ZharfAI Analysis
llama.cpp has found more useful work inside two existing mobile-GPU generations, not announced a faster chip. Release b10687, published at 18:23 UTC on August 29, changes the OpenCL matrix-multiplication route for Qualcomm’s Adreno X2E and A7X families. On an Adreno X2-90, the project says the new default path is worth about 25% of prefill performance for the stock gpt-oss-20b model. On an Adreno 740, bypassing a poorly scheduled tiled kernel is worth about 9% for Gemma 3n E4B. Those are consequential gains for local inference because prefill is the work performed before the first generated token appears. They also came from dispatch and compiler-aware routing rather than replacing silicon.
The mechanism is precise. On the X2-90, llama.cpp’s previous F16-by-F32 kernel ran attention projections at roughly one quarter of the throughput of a tuned dense q4_0 matrix multiply on the same device. Stock gpt-oss-20b spent 40.8% of its prefill GPU time in that single kernel. A faster xmem route already existed but was opt-in; b10687 makes it the default on X2E devices. The Adreno 840 measured neutral, so its behavior stays unchanged. The change also does nothing for the q8-attention variant, which already uses a different dense path. The release reports 963 matrix-multiplication tests passing with none failing across the two test arms.
The A7X fix addresses a different physical constraint. The project says the older compiler allocates 488 bytes of private memory per work item for the tiled F32 kernel, versus 304 bytes for the same source on the next GPU generation. That register pressure causes spilling inside the inner loop. For batched F32-by-F32 work on A7X devices, llama.cpp now falls through to a per-row kernel that the compiler handles better. Small batches keep the tiled route. In both fixes, decode placement is untouched: the relevant gates activate only at larger batch dimensions. Users should therefore expect a shorter wait before output in the tested cases, not a blanket increase in token-generation speed.
That boundary is financially important. A 25% prefill gain can reduce time-to-first-token, increase completed prompts per device-hour, or defer a hardware upgrade if the same improvement holds in a real workload. It does not automatically reduce the cost of every response by 25%. Prefill may be only one share of end-to-end latency; memory, model loading, sampling, power, thermal throttling and idle time remain. The benchmark covers named models, quantization choices and chips, not an average across the installed base. A finance team should treat it as a capacity hypothesis to test with production traces, not as a new depreciation schedule.
Marvell’s Form 10-Q, filed on August 28, shows the market rewarding the physical side of AI capacity at a much larger scale. Revenue for the quarter ended August 1 was $2.739 billion, up 36.5% from $2.006 billion a year earlier. Data-center revenue was $2.172 billion, 79% of the total, versus $1.491 billion and 74% a year earlier; management attributes the 46% data-center increase primarily to strong AI-related demand. Gross margin improved by 2.7 percentage points as higher revenue improved cost absorption and product mix. Net cash from operations for the first six months was $1.2 billion. These figures confirm demand and cash generation for infrastructure components; they do not show how much customer capacity is utilized.
The same filing supplies the counterweight to extrapolation. One unnamed distributor represented 44% of quarterly revenue and one direct customer represented 16%. Sales shipped to customers operating in Asia were about 84% of revenue. Marvell also warns that AI data-center deployments can be delayed by power and water procurement, permitting, supply bottlenecks and local opposition, and that customers may slow or redirect capital spending. Its quarterly research-and-development expense rose 42.8% to $741.1 million. Strong demand therefore arrives with concentration, execution and investment risk rather than as a frictionless margin stream.
Elastic’s August 28 Form 10-Q shows a different place where AI-era spending can create growth without immediate margin expansion. For the quarter ended July 31, total revenue increased 15% to $478.1 million and subscription revenue increased 15% to $448.7 million. Elastic Cloud revenue rose 20% to $235.2 million, while remaining performance obligations reached $1.854 billion, with about 62% expected to be recognized over the next twelve months. Operating cash flow increased to $132.0 million from $104.8 million. Those are real demand and cash signals for enterprise search, observability and security software.
But Elastic’s cost of subscription revenue rose 32% to $91.9 million, more than twice the subscription-revenue growth rate. The company says cloud-hosting costs accounted for $16.0 million of the $22.5 million increase, and subscription gross margin fell to 80% from 82%. Elastic also recorded a $16.7 million net loss and $19.9 million of restructuring and related charges, while holding $1.461 billion in cash, equivalents and marketable securities. The contrast with the llama.cpp patch is useful: software can expose unused performance, but cloud software providers still pay for the underlying capacity. Better kernels help economics only if savings are measurable, retained and not absorbed by higher demand or other infrastructure costs.
The Federal Reserve’s H.8 release adds a credit backdrop without pretending to identify an AI-loan category. Seasonally adjusted loans and leases at U.S. commercial banks increased from $13.9812 trillion in the week ended August 12 to $14.0165 trillion in the week ended August 19, a $35.3 billion rise. Commercial and industrial loans increased $11.8 billion to $2.9457 trillion, while total bank credit rose $33.7 billion to $19.8281 trillion. The weekly levels are estimates subject to benchmarking and seasonal adjustment. They show that aggregate credit was expanding in that week; they cannot tell us whether a data center, chip order or enterprise AI contract received the funds.
The practical test is now straightforward. Device teams should compare time-to-first-token, energy per completed prompt, thermal behavior and total tokens per watt before and after b10687 on the exact model and quantization they deploy. They should separate prefill from decode and report the share of traffic that reaches the new paths. Infrastructure and finance teams should translate only verified production gains into capacity plans, then watch whether usage growth consumes the saving. Investors should distinguish semiconductor revenue, contracted software revenue and general bank credit rather than folding all three into one AI-spending number.
What comes next will decide whether this is a durable efficiency gain or a valuable niche fix. Watch for measurements on more X2E and A7X devices, additional models, longer prompts and sustained thermal loads; regression reports from the default xmem path; and whether similar compiler-aware routing reaches other backends. In company accounts, watch Marvell’s customer concentration, data-center mix and cash conversion, alongside Elastic’s cloud-hosting cost growth and subscription margin. The clearest conclusion today is narrow but actionable: a material slice of AI capacity can be hidden in software choices on hardware already paid for. Finding it is valuable; proving its breadth and keeping the economic benefit are the harder parts.
Sources & documents
- 01Release b10687llama.cpp · August 29, 2026
- 02opencl: use a better matmul path on two Adreno GPU generationsllama.cpp · August 29, 2026
- 03Quarterly report for the quarter ended August 1, 2026Marvell Technology · August 28, 2026
- 04Quarterly report for the quarter ended July 31, 2026U.S. Securities and Exchange Commission · August 28, 2026
- 05Assets and Liabilities of Commercial Banks in the United States — H.8Federal Reserve Board · August 28, 2026
Tags
Related News

llama.cpp Speeds Selected Mac Tests by Fusing Computation
llama.cpp merged selective Apple-GPU optimizations on September 19, with a contributor-reported generation gain of about 16% in one M2 Ultra test. Another attempted fusion was discarded after slowing execution, making workload-specific evidence more useful than a blanket speed claim. A separate CUDA correction addresses a memory error; neither change establishes universal gains or lower operating costs. In the separate macroeconomic picture, August US industrial output was flat and euro-area consumers' one-year inflation expectations rose slightly. Virginia's new executive order also restricts access to state assistance for certain new data-center projects.

Accenture Agrees to Embed Safety Evaluators at Anthropic
Accenture and Anthropic have agreed to establish an embedded evaluation team, with each company expecting at least $1 billion in AI-safety investment over five years. That is a plan, not completed spending or demonstrated safety. California has set deadlines for independent-oversight recommendations, while Michelle Bowman's banking remarks expose the distinction between recognizing risk and acting on it. Separately, Japan's new policy-rate target takes effect on September 24. ZharfAI examines why evaluator access, decision authority and evidence of corrective action should be assessed separately, without treating these different developments as one causal story.

Google and NVIDIA Seek Faster Grid Connections Through Flexible Data Center Power Use
Google, NVIDIA and Emerald AI's new alliance proposes a practical exchange: more flexible electricity demand for faster data center connections. It has not announced newly delivered power or binding rules. ZharfAI examines what that promise would require in an AI service contract and why lower grid draw is not necessarily lower total energy use. The House's separate 417–3 vote on ratepayer protection puts infrastructure-cost allocation in focus, while the Federal Reserve's quarter-point rate increase adds a distinct financing consideration. The tests ahead are regulatory decisions, measured performance and clear responsibility for costs—not the number of companies supporting an announcement.