If you landed here from a search engine, this is probably what you are seeing:
- A guest hangs while printing to a parallel port (
LPT1), usually part-way through a job, and never comes back. - QEMU itself stops responding. The HMP monitor accepts a connection but never answers
info status, and the display stops redrawing. SIGTERMappears to be ignored and onlykill -9ends the process.- It happens with a
socketchardev whose peer is connected but has stopped reading — a print spooler that stalled, for instance. Afilechardev never triggers it. - Small jobs do not reproduce it. A one-megabyte job can finish normally while nothing on the far end reads a single byte.
1. What the VM is doing when it stops
Every time the guest raises the STROBE line, hw/char/parallel.c hands one byte to the character device backend:
if ((s->control & PARA_CTR_STROBE) == 0)
qemu_chr_fe_write_all(&s->chr, &s->dataw, 1);
That call does not return until the byte is gone. In chardev/char.c it becomes a retry loop:
retry:
res = cc->chr_write(s, buf + *offset, len - *offset);
if (res < 0 && errno == EAGAIN && write_all) {
g_usleep(100);
goto retry;
}
The caller is a vCPU thread, and it holds the BQL. A stack taken while the VM was stuck shows both halves of the problem at once:
TID (vCPU) TID (main thread)
g_usleep __lll_lock_wait
qemu_chr_write_buffer __pthread_mutex_lock
qemu_chr_write qemu_mutex_lock_impl
parallel_ioport_write_sw main_loop_wait
memory_region_write_accessor
The left column starves the right one. The guest cannot advance because its vCPU is inside the loop, and the main loop cannot run because it never gets the lock — which is why the monitor goes silent and the display stops.
SIGTERM is not being ignored. It is waiting its turn behind the stuck write. In one run the signal was delivered, nothing happened for twelve seconds, and the moment the far end resumed reading, QEMU finished the write and exited immediately — processing the signal that had been queued all along.2. Why the code is written that way
The comment above that call has been in the tree since 2016, in commit 6ab3fc32ea:
/* XXX this blocks entire thread. Rewrite to use
* qemu_chr_fe_write and background I/O callbacks */
That commit is worth reading, because it moved away from the non-blocking call:
- qemu_chr_fe_write(s->chr, &s->dataw, 1);
+ /* XXX this blocks entire thread. Rewrite to use
+ * qemu_chr_fe_write and background I/O callbacks */
+ qemu_chr_fe_write_all(s->chr, &s->dataw, 1);
Its message explains why: almost no caller of qemu_chr_fe_write() checked the return value, so an EAGAIN meant bytes were silently dropped. Data loss was traded for a blocked thread, and the comment records the cost of that trade. The fix below pays neither price.
3. Small jobs will not reproduce it
The kernel socket send buffer absorbs the job before the write ever sees EAGAIN. If the job is smaller than that buffer, everything completes normally even though the peer never reads a byte.
| Configuration | Stalls after |
|---|---|
| Linux harness, sink with a 4 KB receive buffer | 1,357,824 bytes |
| Windows XP guest | 1,792,000 bytes |
| For scale — one Windows XP printer test page | 146,613 bytes |
On XP, copy /b c:\windows\explorer.exe lpt1: (1,025,049 bytes) completes normally against a sink that reads nothing at all. A printer test page is a twelfth of the threshold. So a stalled backend plus a small job looks perfectly healthy — do not conclude from that test that you are safe.
4. The fix: assert BUSY instead of blocking
A real printer that cannot accept data keeps its BUSY line asserted, and the host side of a parallel port has no way to stall the bus. So when the backend is full, do the same thing:
- Use the non-blocking
qemu_chr_fe_write(). OnEAGAIN, arm aG_IO_OUTwatch and return immediately. - Leave BUSY asserted while the write is outstanding. The guest driver waits the way it would for a slow printer, and the vCPU thread is free.
- Suppress the ACK/BUSY restore sequence in the status-register read while a write is pending — otherwise the guest pushes the next byte into a full pipe.
- Withhold the acknowledge interrupt until the byte has really gone out.
- If a strobe arrives anyway, the guest ignored BUSY; drop that byte, as real hardware would garble it.
- The pending byte is migration state, so it travels in a vmstate subsection and
post_loadre-arms the watch.
A file chardev never returns EAGAIN and is unaffected. An unconnected socket still fails with EIO and drops the byte, exactly as before.
5. What changed, measured
Same reproducer, same stalling sink, two builds differing by this patch alone:
| Before | After | |
|---|---|---|
HMP info status while stalled | no answer | VM status: running |
| vCPU thread state | hrtimer_nanosleep | kvm_vcpu_block |
| Main thread state | futex_do_wait | poll_schedule_timeout |
SIGTERM | ignored 10-12 s, needs kill -9 | exits within 2 s |
| Bytes delivered once the peer resumes reading | not guaranteed | all of them, byte-exact |
The guest stalls at the same byte in both builds — 1,792,000 on XP — which is the point. The patch does not remove the back-pressure; it moves it to the layer that is supposed to carry it. Before, the host froze. After, only the guest waits.
Nothing is lost. A Linux guest pushed 1,638,400 bytes (200 × 8192) and a Windows XP guest 4,126,720 bytes (4 × 1,031,680) through a sink that stalled for three minutes; both arrived complete and byte-exact once it resumed reading. With the interrupt-driven path exercised, the parallel port interrupt count went from 0 to 1,638,600, so the deferred-acknowledge path is covered too.
6. Reproducing it
The harness is a Linux guest built from an 8 MB kernel and a 1 MB static busybox. It boots to a result in 0.2 seconds and touches the parallel port zero times when idle, which is the whole point — on a desktop guest, X and udev generate enough register traffic that you cannot attribute anything.
guest/fetch-deps.sh # kernel, busybox, parport modules
guest/build-initramfs.sh
export LPT_BIN=/path/to/qemu-system-i386
scripts/stall-sink.py 600 /tmp/sink.bin &
NOTRACE=1 FLOODN=1024 \
LPT_CHARDEV="socket,id=lptfile,host=127.0.0.1,port=9100,server=off" \
TMO=240 scripts/lpt-linux.sh freeze flood
To tell a frozen host from a guest that is merely busy-waiting, ask the monitor. A blocking write takes the BQL down with it; a guest spinning on its own does not.
echo info status | socat -T5 - UNIX-CONNECT:logs/freeze-mon.sock
A related failure that is not this one
While chasing this, a Windows XP guest was also seen spinning on the status register — 38,433 reads, every one returning 0xd8, never changing. That is an IEEE-1284 negotiation with nothing on the other end to answer it, and it is a different failure: the guest busy-waits, but QEMU stays healthy and the monitor answers normally. It also never reaches the blocking write, because STROBE is never raised during that sequence. Two symptoms that look alike from inside the guest, with different causes.
Not submitted upstream
QEMU declines contributions that include or derive from AI-generated content, and the exception list is empty. This investigation and patch were produced with Claude, so the patch route is closed. The same policy explicitly does not cover using AI for debugging, so a bug report would have been possible; it was not filed.
This patch is therefore not in QEMU. If you want it, you have to apply it yourself, and the verification above is all there is.
Tested on
QEMU 11.1.50 (v11.1.0-1474-g43922bf55e) on Ubuntu 26.04 with KVM, against a q35 Windows XP guest and a minimal Linux guest on the Alpine 6.12 x86 kernel. Full write-ups, traces and the captured stack are in the repository (detailed documents are in Korean).