HN
Paper
All
Show
Ask
Jobs
Top
Today
Last 7 days
Last months
This year
Statistics
All
Show
Ask
Jobs
Top stories
Today
Last 7 days
Last months
This year
Statistics
Stories by
somnial
The case for disaggregated LLM serving
4 points
somnial
2026-08-12T08:46:51Z
blog.doubleword.ai
On-the-fly snapshot compression for elastic inference at scale
6 points
somnial
2026-08-11T12:40:12Z
blog.doubleword.ai
NVLink, NVSwitch, and All That
5 points
somnial
2026-07-22T10:59:11Z
blog.doubleword.ai
The Anatomy of an Instruction Pipeline Hazard
9 points
somnial
2026-07-14T07:06:50Z
hiraditya.github.io
Width vs. Depth: Speculating on the Margin
17 points
somnial
2026-07-02T14:52:31Z
blog.doubleword.ai
Pushing memory bound CUDA kernels past the speed of light with data compression
2 points
somnial
2026-05-28T12:42:26Z
fergusfinn.com
Speculative KV coding: ~4× losslessly compressed KV cache using a small model
2 points
somnial
2026-05-12T13:18:37Z
fergusfinn.com
70x faster cold(ish) starts for SGLang
1 points
somnial
2026-04-27T16:18:05Z
fergusfinn.com
LLM powered data structures: A lock-free binary search tree
1 points
somnial
2026-01-13T09:57:31Z
fergusfinn.com
Parallel Primitives for Multi-Agent Workflows
1 points
somnial
2026-01-05T18:54:34Z
fergusfinn.com
Scheduling in LLM Inference
1 points
somnial
2025-11-14T07:39:35Z
fergusfinn.com
How fast can an LLM go?
2 points
somnial
2025-10-30T10:18:32Z
fergusfinn.com