Humming: a vllm kernel library for running glm-5.3-flash on sparks
An inference kernel library for running GLM-5.3-Flash on DGX Sparks.
Problem it solvesLLM inference on Spark hardware needed support in vLLM.
Joseph Sauvage reports that Humming, a new kernel library inside vLLM, now runs his uncensored GLM-5.3-Flash build on four DGX Sparks. He says the setup uses Humming, FlashInfer, and his own quantization and patches, with reported gains in speed and cache hit rate.
Joseph Sauvage
@JoesInvestments
Humming season has begun! Marlin has run 4-bit inference for years. This week my uncensored GLM-5.3-Flash on four DGX Sparks moved its experts off it. They now run on Humming, the new kernel library inside vLLM, which added support for the Spark's chip six days ago. As far as I can find, this is the first public GLM-5.3-Flash build on it. None of it exists without these people. Z.ai (@Zai_org) open sourced GLM-5.3-Flash. Blackfrost (@Blackfrost_AI) made it uncensored. Z Lab (@zhijianliu_) created DFlash, and incoai trained the DFlash2 drafter that makes this model fast (huggingface.co/incoai/GLM-5.3…). Tony (@2WildTech) got it running on four Sparks first. Eight of my patches are ports of his GB10 work, and both FlashInfer fixes are his. Jacopo Nardiello (@jnardiello) proved the 8-bit drafter on his own Spark build. My port of it made agent sessions 8.6% faster. Jinzhen Lin and Julian Huang at Ant Group built Humming, with @mgoin_ among its contributors (github.com/vllm-project/h…). The @vllm_project team makes the engine under all of it, and LLM Compressor, the tool I quantized with. FlashInfer runs the attention. Matt Mastracci (@mmastrac) and Jared Wen fixed three bugs that broke how this model reads anything past 2,048 tokens. Juntian Liu, weijie and Harshil wrote the other vLLM fixes I carry. Stanislav Bardyuk wrote the NCCL fix for the network deadlock I reported, and Rami Nudelman at NVIDIA brought it to me. Onur Solmaz (@onusoz) wrote oomwrap, which guards every launch. Here is what I added. I quantized the whole model myself. 643 GB of Blackfrost's full-precision weights, down to 4-bit on my own four Sparks, with a scale search I picked after testing the options on real weights. It has 2% lower error than the conversion I shipped three days ago, and the same agent tasks finish 20% sooner because fewer answers run out of room. Four of the 18 patches in the build are mine, two of them ports of Jacopo's experiments. So is every build script, launcher and benchmark in the repo. Resumed agent sessions now keep their cache instead of recomputing it, so the first token comes about a third sooner on long sessions. And the model always thinks at max, the way my agents use it. I also wrote short-turn caching into vLLM's cache manager, with upstream fixes from Netanel Haber and Nick Hill. Short agent turns went from 0% to 59% cache hits. It failed one correctness check I set before the run, so it waits. Three days ago my build ran on Tony's conversion recipe and the vLLM 0.30 release. Today it runs my own quantization on vLLM's development branch and its newest kernels. I learned a ton getting here. The newer path costs about 7% throughput on long agent sessions against my 0.30 build. I took that trade. Fixes land on main first, and Humming is young, with room to grow. Want to try it? With four Sparks, the README takes you from a stock vLLM image to a running server: one build script and one launch command per node. The build checks every patched file byte for byte against the image I serve. One user gets 38 to 128 tok/s depending on what it writes. 32 users get 572 tok/s together. Uncensored, always reasoning. *Open doors* Where vLLM main loses 5% against 0.30. I have not profiled it yet. The drift that keeps short-turn caching out: one resume in eight. Noise or bug. Repo: github.com/joesinvestment…


