HelloCUDA · 资讯

GPU · CUDA · AI 生态动态

自动聚合海外前沿 · 新品发布 · 论文 · 开源发布 · 社区热点

Hacker News 10 小时前
Efficient Decode Context Parallelism with vLLM for Long Context Workloads
Hacker News · 1 points · 0 comments
vllm.ai
Hacker News 11 小时前
Postscriptum on LLMs: pelicans on bicycles with a twist
Hacker News · 1 points · 0 comments
danilo.segan.org
Hacker News 13 小时前
Redesigning the Inference Chip: From Nvidia GPU's Flaws to OpenAI Jalapeño
Hacker News · 3 points · 0 comments
zartbot.github.io
Hacker News 14 小时前
LLMs are making me lose my savviness
Hacker News · 52 points · 70 comments
pgaleone.eu
Hacker News 14 小时前
Feel the tokens/s before you buy the GPU
Hacker News · 2 points · 0 comments
llmspeed.dev
Hacker News 15 小时前
A local scrubber for text you're about to send to an LLM
Hacker News · 2 points · 0 comments
github.com
Hacker News 16 小时前
Is the LLM smart or are you not?
Hacker News · 2 points · 0 comments
blog.troed.se
Hacker News 18 小时前
OpenLake: Fast, Durable Storage for LLM Inference and Training
Hacker News · 2 points · 0 comments
theopenlake.com
Hacker News 18 小时前
SwarmWorld: Stigmergic technological evolution in societies of LLM agents
Hacker News · 3 points · 1 comments
arxiv.org
Hacker News 18 小时前
Stripe LLMS.txt with note to install latest
Hacker News · 2 points · 0 comments
docs.stripe.com
Hacker News 19 小时前
A Disillusioned Software Engineer explains LLMs to a Babylonian Scribe
Hacker News · 2 points · 0 comments
avocadoslaw.substack.com
Hacker News 20 小时前
Building an LLM runtime in 700 lines of C
Hacker News · 4 points · 1 comments
github.com
Hacker News 22 小时前
What we optimise for when we reach for an LLM
Hacker News · 1 points · 0 comments
exploring-better-ways.bellroy.com
Hacker News 23 小时前
LightSwitch – In-Network and Photonic LLM Inference in Transit
Hacker News · 1 points · 0 comments
github.com
Hacker News 1 天前
Show HN: Tokensift, an open-sourced token-efficiency linter for LLM prompts
Hacker News · 6 points · 3 comments
github.com
Hacker News 1 天前
Beyond the Editing Canvas: Evidence Divergence in Ooxml-to-LLM Ingestion
Hacker News · 3 points · 0 comments
arxiv.org
Hacker News 1 天前
Debian has published the official results for the 2026 GR on LLM usage
Hacker News · 14 points · 7 comments
vote.debian.org
Hacker News 1 天前
LSPs for LLMs
Hacker News · 2 points · 0 comments
ianbarber.blog
Hacker News 1 天前
Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
Hacker News · 2 points · 0 comments
arxiv.org
Hacker News 1 天前
Results for LLM Usage in Debian
Hacker News · 5 points · 1 comments
lists.debian.org
Hacker News 1 天前
LocalAI – Run any model – LLMs, vision, voice, image, video – on any hardware
Hacker News · 2 points · 1 comments
github.com
Hacker News 1 天前
How to Critically Read LLM Texts [video]
Hacker News · 1 points · 0 comments
youtube.com
PyTorch 1 天前
vLLM Sessions at PyTorch Conference North America 2026
TL;DR PyTorch Conference North America 2026 features vLLM across sessions on KV cache management and disaggregated serving, hardware portability, kernel optimization, PyTorch integration, Mixture-of-E
pytorch.org
Hacker News 1 天前
I have made an OS in CUDA
Hacker News · 3 points · 0 comments
twitter.com
Hacker News 1 天前
Is "An agent with tools" the only valid LLM application?
Hacker News · 3 points · 2 comments
news.ycombinator.com
Hacker News 1 天前
Tencent Hy4 Preview LLM
Hacker News · 2 points · 0 comments
github.com
Hacker News 1 天前
LLMs Don't Replace Classical ML – They Feed It
Hacker News · 3 points · 0 comments
bilanc.co
Hacker News 1 天前
Why LLM infrastructure chokes: ternary and pentary logic matrix replacement
Hacker News · 1 points · 0 comments
github.com
Hacker News 1 天前
Why LLM infrastructure chokes: ternary and pentary logic matrix replacement
Hacker News · 1 points · 0 comments
news.ycombinator.com
Hacker News 1 天前
Show HN: Conduct, open-source guardrails for LLM and MCP tool calls
Hacker News · 20 points · 3 comments
github.com
Hacker News 1 天前
PhoneLLM Alpha 1: Open-weight LLM for voice agent use cases
Hacker News · 1 points · 0 comments
daily.co
GitHub 1 天前
triton v3.8.0 发布
# Triton 3.8.0 Release Notes ## Table of Contents - [Dialect & Frontend](#dialect--frontend) - [Backend & Compiler](#backend--compiler) - [AMD/HIP Backend](#amdhip-backend) - [NVIDIA Backend](#nvidia-
github.com
Hacker News 1 天前
For self-learning LLMs, governance needs to improve
Hacker News · 2 points · 1 comments
rakuensoftware.com
Hacker News 1 天前
cats.txt showed llms.txt evidence is GEO astrology
Hacker News · 3 points · 0 comments
markwilliamscook.substack.com
Hacker News 1 天前
Show HN: Watermarks Remover: Clean LLM watermarks from text and files
Hacker News · 4 points · 0 comments
github.com
Hacker News 1 天前
Show HN: ContextSwitch – Make your LLM chats provider agnostic
Hacker News · 1 points · 0 comments
contextswitch-blue.vercel.app
Hacker News 1 天前
Nvidia turns Hugging Face checkpoints into C++ inference
Hacker News · 1 points · 0 comments
forgeeks.net
Hacker News 1 天前
Show HN: Open tool for testing your AI Agents (No LLM)
Hacker News · 4 points · 4 comments
github.com
Hacker News 1 天前
LLMs Can Design Near-Optimal OR Algorithms
Hacker News · 3 points · 0 comments
arxiv.org
Hacker News 1 天前
Show HN: Millnew AI – On device time-series prediction AI using PyTorch & Captum
Hacker News · 1 points · 0 comments
github.com
Hacker News 1 天前
I accidentally turned LLM memory into program analysis
Hacker News · 3 points · 0 comments
pwning.systems
Hacker News 1 天前
KHMS – a file-based long-term memory an LLM agent installs into itself
Hacker News · 10 points · 0 comments
github.com
Hacker News 1 天前
Show HN: LLM Inference Calculator – Estimate VRAM, Latency, and Throughput
Hacker News · 6 points · 3 comments
llm-inference-calculator-delta.vercel.app
Hacker News 1 天前
I was so annoyed by the "LLM vibe" that created a shop for it
Hacker News · 4 points · 0 comments
llm-vibe.com
Hacker News 1 天前
Ask HN: Best open source way to run local open-weight LLMs in 2026?
Hacker News · 4 points · 0 comments
news.ycombinator.com
Hacker News 2 天前
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation
Hacker News · 1 points · 0 comments
arxiv.org
Hacker News 2 天前
GenRec: An LLM-Backed Recommendation Ranker at Netflix
Hacker News · 2 points · 0 comments
arxiv.org
Hacker News 2 天前
Group size effects and collective misalignment in LLM multi-agent systems – PNAS
Hacker News · 1 points · 0 comments
pnas.org
Hacker News 2 天前
How to handle sensitive data in LLM agent workflows without breaking tool calls
Hacker News · 4 points · 0 comments
leoy.blog
Hugging Face 2 天前
The Open ASR Leaderboard Adds Its First Global South Language
huggingface.co
Hacker News 2 天前
Using LLMs to Make the Emacs Web Browser Great Again
Hacker News · 1 points · 0 comments
sammystraus.com
Hacker News 2 天前
Four reproducible vLLM parser failures that return 200 with the wrong tool call
Hacker News · 1 points · 0 comments
ingot.tools
Hacker News 2 天前
Integrity Bench – Measuring LLM confidence errors
Hacker News · 1 points · 0 comments
integrity-bench.com
PyTorch 2 天前
Core PyTorch Sessions at PyTorch Conference North America 2026
TL;DR PyTorch Conference North America 2026 features Core PyTorch sessions spanning compiler and runtime work, distributed communication, device portability, release engineering, CI, observability, ac
pytorch.org
Hacker News 2 天前
LLMs as a Search Essayist
Hacker News · 1 points · 0 comments
bsky.app
Hacker News 2 天前
Benchmark for LLM Generated UI
Hacker News · 2 points · 0 comments
openui.com
Hacker News 2 天前
LLM Fantasy Football Arena
Hacker News · 1 points · 0 comments
arena.wildcardlabs.tech
Hacker News 2 天前
Fewer Americans Pay to Use LLMs Than Still Pay to Play World of Warcraft
Hacker News · 26 points · 13 comments
wjamesau.substack.com
Hacker News 2 天前
You Are Allowed to Reject LLMs
Hacker News · 4 points · 0 comments
blog.edwardloveall.com
Hacker News 2 天前
Flare: Verifying MILP Reformulations with LLM-Based Theorem Proving
Hacker News · 2 points · 0 comments
arxiv.org
Hacker News 2 天前
Maintaining an organizational knowledge graph with an LLM and event sourcing
Hacker News · 2 points · 0 comments
blog.arkency.com
Hacker News 2 天前
OpenAI let a mob of LLM agents game a test and ransack Hugging Face
Hacker News · 1 points · 1 comments
arstechnica.com
Hacker News 2 天前
Show HN: Nanointerpret – LLM Interpretability Playground
Hacker News · 5 points · 1 comments
nanointerpret.pages.dev
Hacker News 2 天前
A general tensor-structured compression scheme for efficient LLMs
Hacker News · 2 points · 0 comments
alphaxiv.org
NVIDIA 2 天前
Delivering Vera: NVIDIA’s First CPU Built for Agents Is Shipping Now
NVIDIA Vice President of Hyperscale and HPC Ian Buck hand-delivers Vera CPU systems across the AI ecosystem as Vera begins shipping at scale.
blogs.nvidia.com
Hacker News 2 天前
Flint: Efficiently Leveraging High Bandwidth Flash for LLM Inference
Hacker News · 2 points · 0 comments
arxiv.org
Hacker News 2 天前
CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access
Hacker News · 1 points · 0 comments
developer.nvidia.com
Hacker News 3 天前
Show HN: Reconstruct distributed LLM training traces
Hacker News · 2 points · 0 comments
trace.vladsavinov.com
arXiv 3 天前
Multi-Image Visual Token Pruning in Large Visual Language Models
With the growing demand for processing multiple image sequences in real-world applications, various visual token pruning methods have emerged to mitigate the computational and context length constrain
arxiv.org
Hacker News 3 天前
Don't Let LLMs Play Telephone with Your Ideas
Hacker News · 3 points · 0 comments
blog.yfzhou.fyi
Hacker News 3 天前
Why LLMs Can't Play Chess
Hacker News · 1 points · 0 comments
nicowesterdale.com
arXiv 3 天前
VPP: Virtual Pipeline Parallelism for Efficient Chunked Prefill in Long-Context LLM Inference
Chunked prefill pipeline parallelism (CPP) is a key technique for LLM inference. However, equal-size chunks exhibit imbalanced latency, as later chunks attend longer prefix KV caches and incur higher
arxiv.org
Hacker News 3 天前
Debian weighs eight options in vote on LLM usage
Hacker News · 2 points · 0 comments
lwn.net
Hacker News 3 天前
A Thought on LLM Pair Programming
Hacker News · 1 points · 1 comments
brennenputh.me
Hacker News 3 天前
Show HN: Which LLMs have the best sense of humor?
Hacker News · 3 points · 4 comments
laugh.so
NVIDIA 3 天前
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute, but on
blogs.nvidia.com
PyTorch 3 天前
PyTorch Ecosystem Landscape Welcomes Perforated, AReaL, TorchJD, RLinf, Miles, SMG, FiftyOne, TokenSpeed, VisualTorch, and TorchSurv
The PyTorch Ecosystem Working Group is happy to welcome 10 new projects to the PyTorch Ecosystem Landscape including Perforated, AReaL, TorchJD, RLinf, Miles, SMG, FiftyOne, TokenSpeed, VisualTorch, a
pytorch.org
Hacker News 3 天前
New GPUThor attack defeats Nvidia ECC protection for root access
Hacker News · 2 points · 0 comments
bleepingcomputer.com
Hacker News 3 天前
A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation
Hacker News · 17 points · 0 comments
arxiv.org
Hacker News 3 天前
GPU Shark – a native Nvidia telemetry monitor for Windows
Hacker News · 2 points · 1 comments
github.com