blockjob: Lie better in child_job_drained_poll()

Block jobs claim in .drained_poll() that they are in a quiescent state as soon as job->deferred_to_main_loop is true. This is obviously wrong, they still have a completion BH to run. We only get away with this because commit 91af091f92 added an unconditional aio_poll(false) to the drain functions, but this is bypassing the regular drain mechanisms. However, just removing this and telling that the job is still active doesn't work either: The completion callbacks themselves call drain functions (directly, or indirectly with bdrv_reopen), so they would deadlock then. As a better lie, tell that the job is active as long as the BH is pending, but falsely call it quiescent from the point in the BH when the completion callback is called. At this point, nested drain calls won't deadlock because they ignore the job, and outer drains will wait for the job to really reach a quiescent state because the callback is already running. Signed-off-by: Kevin Wolf <kwolf@redhat.com> Reviewed-by: Max Reitz <mreitz@redhat.com>
2025-08-05 16:53:55 -06:00 · 2018-09-07 15:31:22 +02:00 · 2018-09-07 15:31:22 +02:00 · b5a7a05735
commit b5a7a05735
parent 46aaf2a566
3 changed files with 14 additions and 2 deletions
--- a/job.c
+++ b/job.c
@ -857,7 +857,16 @@ static void job_exit(void *opaque)
    AioContext *ctx = job->aio_context;

    aio_context_acquire(ctx);
+
+    /* This is a lie, we're not quiescent, but still doing the completion
+     * callbacks. However, completion callbacks tend to involve operations that
+     * drain block nodes, and if .drained_poll still returned true, we would
+     * deadlock. */
+    job->busy = false;
+    job_event_idle(job);
+
    job_completed(job);
+
    aio_context_release(ctx);
 }

@ -872,8 +881,8 @@ static void coroutine_fn job_co_entry(void *opaque)
    assert(job && job->driver && job->driver->run);
    job_pause_point(job);
    job->ret = job->driver->run(job, &job->err);
-    job_event_idle(job);
    job->deferred_to_main_loop = true;
+    job->busy = true;
    aio_bh_schedule_oneshot(qemu_get_aio_context(), job_exit, job);
 }