Tech • AI • Robotics • Game

VIDEO
ENFR

Full article — scored 10/10

Decode-Latency Feedback Prefill: Model-Free Controller for Autoregressive Inference

A new arXiv paper introduces Decode-Latency Feedback Prefill, a scheduler-side controller for autoregressive inference that aims to smooth decode latency by resizing prefill chunks from recent latency feedback. The result is promising but deliberately bounded: it improves P99 inter-token latency for one small-model, single-GPU setup, while failing to generalize to larger models and multi-GPU execution.

Sign in to follow
Generated October 1, 2026 at 6:53 AM1680 wordsOriginal source — ArXiv - Artificial Intelligence

A narrow but important latency problem

Decode-Latency Feedback Prefill, or DLFP, is a fresh systems proposal for a familiar serving pain point: in autoregressive model inference, a newly admitted prompt can interrupt the rhythm of tokens already being generated for other users or tasks . The paper, submitted to arXiv on September 29, 2026, frames the issue as an interference problem between the prefill phase, where the model processes prompt tokens and builds state, and the decode phase, where it emits output tokens one step at a time .

The headline matters because the proposed fix is not a new transformer architecture, a new attention kernel, or a model-specific latency predictor. DLFP is a model-free controller placed in the scheduling path: it observes recent scheduling intervals and adjusts how much prefill work may overlap with active decodes . In other words, it treats latency as a feedback signal and tries to regulate prefill admission before that prefill work causes visible pauses in streamed output.

The current public record is also unusually careful about limits. The authors report a positive result for Qwen3-0.6B in BF16 on a single NVIDIA A100 80 GB GPU, but they explicitly say the same mechanism does not generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel configuration . That makes DLFP less a finished serving feature than a research prototype with a clear lesson: feedback can help, but only when the feedback signal actually tracks completed device work.

What DLFP changes — and what it leaves alone

Autoregressive serving commonly alternates between two different workloads. Prefill processes the input prompt and exposes parallelism across the prompt tokens; decode generates one token per active sequence per iteration and is sensitive to anything added to the critical path . In continuous batching, a long prompt admitted next to requests that are already decoding can make the next mixed iteration longer, so multiple active decoders may experience the same latency spike .

DLFP targets only the portion of prefill that overlaps active decode work. If no request is decoding, the controller does not constrain isolated prefill; if at least one request is decoding, it caps the prefill chunk size using a proportional update based on the observed interval from a guarded scheduling cycle . The evaluated policy sets a target mixed-iteration latency of 80 milliseconds, starts with a 3,072-token cap, clamps the cap between 1,024 and 6,144 tokens, and quantizes updates to 128-token increments .

That design choice is the central claim. Fixed chunked prefill can reduce interference, but a fixed cap must be tuned for a specific combination of model, hardware, workload, and latency objective . DLFP instead asks whether recent measured delay can stand in for a hand-built or learned latency model. If the observed interval is above target, the next chunk shrinks; if it is below target, the next chunk grows .

The implementation described in the paper modifies the vLLM V1 scheduler while leaving model weights, BF16 precision, attention backend, sampling, KV layout, and isolated-prefill scheduling unchanged . That is why the work is best read as a serving-scheduler experiment rather than a model-architecture proposal.

The positive result: smoother token cadence on one small-model setup

The main experiment is tightly specified. The authors run Qwen3-0.6B in BF16 on one NVIDIA A100 80 GB GPU with vLLM 0.17.1, PyTorch 2.10, CUDA 12.8 user-space libraries, FlashAttention-2, CUDA graphs, and asynchronous scheduling . The primary workload uses nonmatching Agent prompts at an offered load of 1.0 request per second, with 100 measured requests after warmups and paired seeds 801, 802, and 803 .

In that setup, DLFP reduces P99 inter-token latency by 24.8%, 30.1%, and 28.2% across the three paired trials, for a mean reduction of 27.7% and a paired 95% confidence interval of 21.0% to 34.3% . The authors also report exact output agreement, no failures, and unchanged SLO compliance in the positive trials .

Those details are important because latency optimizations can easily hide quality, correctness, or reliability regressions. Here, the paper says all 600 baseline and candidate requests completed successfully, workload fingerprints matched, and all 76,800 paired generated token IDs agreed exactly . For a scheduling-side change, exact-token agreement is a strong sanity check: DLFP did not get its latency gain by altering decoding decisions.

The cost is equally explicit. Mean P99 time to first token rises by 34.8% while remaining inside the declared SLO, and mean P99 end-to-end latency rises by 7.8% with a confidence interval that crosses zero . In practical terms, DLFP improves the smoothness of streamed tokens after generation begins, but it can make users wait longer for the first token. The paper therefore positions the controller as useful only when interactive token cadence is more important than minimizing time to first token .

The negative result is not a footnote

The most valuable part of the paper may be its refusal to generalize the positive result. In a two-GPU Qwen3-0.6B tensor-parallel experiment, the authors report paired P99 inter-token-latency reductions of 5.3%, 8.2%, and 24.7%, with a mean confidence interval crossing zero; SLO passes fall, P99 time to first token rises, and energy per output token increases . The paper calls this “not a win” .

The larger-model results are harsher. For Qwen3-8B at 0.10 requests per second, DLFP raises P99 inter-token latency from 19 milliseconds to 281 milliseconds and reduces SLO passes from 10 out of 20 to 7 out of 20 . At 0.25 requests per second, it improves P99 inter-token latency from 1.763 seconds to 707 milliseconds, but P99 time to first token rises and SLO passes drop to zero, so the controller moves latency rather than restoring useful goodput .

For Qwen3-32B on two GPUs, the paper reports worse P99 inter-token latency and a 4.7% energy-per-output-token regression, while a completion-aware revision also fails under synchronous scheduling . That makes the negative result structural rather than a simple matter of retuning a cap.

The root cause, according to the authors, is that the small-model success depends on an accidental correlation between scheduler-call cadence and completed GPU iteration time . Larger kernels, asynchronous queueing, and tensor-parallel synchronization break that correlation, so the controller can expand its prefill cap after seeing a short host-side interval even when device work is still outstanding . A feedback controller is only as trustworthy as the signal it observes.

Why “model-free” is attractive — and dangerous

The appeal of DLFP is clear. Analytical or learned latency models must account for model size, attention shapes, batch composition, kernels, device generation, clock state, and competing work . That is a heavy burden, especially for heterogeneous laptops, phones, and local inference environments where a deployment may not look like a controlled datacenter GPU benchmark .

DLFP tries to avoid that burden by closing the loop directly. Instead of predicting how expensive the next mixed iteration will be, it observes recent delay and adjusts the next prefill allowance . This is attractive because it could adapt to changing load without explicit knowledge of the model architecture or hardware profile.

But the paper’s generalization failures show the danger. “Model-free” does not mean “measurement-free.” If the observed interval is merely a host-side proxy and not the actual completed GPU iteration time, the controller can confidently act on the wrong signal . The reported failures therefore define the boundary of the contribution: DLFP is evidence that feedback prefill control can work at one operating point, not evidence that the current controller is ready for broad deployment.

Edge inference is a motivation, not a proven result

The subject has an obvious edge-computing angle. The paper describes scenarios such as a laptop assistant generating an answer while indexing a document, or a phone running an interactive agent beside background summarization . Small open-weight models make that kind of local concurrency plausible, but edge devices introduce responsiveness, power, and thermal constraints .

Still, the authors are explicit that they do not claim mobile-device performance . They also note that a single interactive phone session has no overlapping prefill and decode, so DLFP cannot help unless multiple local tasks overlap . The next target, in their view, should be a completion-timed controller evaluated on real concurrent CPU and mobile workloads, with device-completion events or runtime callbacks rather than scheduler-call cadence .

That distinction matters for product readers. DLFP should not be marketed as a phone-speedup result. It is better understood as a proof-of-concept that points toward completion-aware adaptive prefill for local and heterogeneous inference, provided future experiments measure P99 inter-token latency, time to first token, SLO goodput, energy per token, quality, and sustained thermal behavior .

What changes now

The current state of DLFP is therefore precise. It is a newly published arXiv v1 paper, not a production feature or a generally validated serving policy . It demonstrates a meaningful P99 inter-token-latency reduction in one carefully controlled Qwen3-0.6B single-GPU setup, while documenting that the same feedback mechanism breaks on larger and multi-GPU configurations .

That balance is the story. The field does not get a universal prefill controller this week. It gets a compact experiment showing that decode-latency feedback can be useful, plus a warning that host-side scheduling intervals are a fragile control signal under asynchronous GPU execution. For inference engineers, the actionable insight is not simply “use DLFP.” It is: protect decode latency, measure the right completion signal, and treat prefill chunking as a closed-loop control problem only when the loop is connected to the hardware reality it is trying to regulate.

Sources from the last 72 hours

  1. [1]Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization LimitsSep 29, 2026, 8:44 PM
  2. [2]Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization LimitsSep 29, 2026, 8:44 PM
  3. [3]Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization LimitsSep 29, 2026, 8:44 PM

AI-generated article based on recent web research, then preserved as a dated editorial snapshot.