Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 114 words in the discussion.

Trending Topics

#vllm#don#model#tech#writes#https#github#com#pull#llama

Discussion (6 Comments)Read Original on HackerNews

ilc•about 1 hour ago
Watch the video carefully. DFlash2's tool call fails on python syntax.

Usually models in this class nail things like that 1 shot, which the other side did.

I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.

hypfer•about 3 hours ago
Amazing tech

> An agent writes in an afternoon what a chatbot writes in a month

But can you just.. not.

Your tech is so good, it speaks for itself. Don't ruin that.

adefa•about 3 hours ago
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
sarjann•about 2 hours ago
Great news, has made low memory bandwidth model usage so much nicer.
verdverm•about 3 hours ago
cogman10•about 2 hours ago