Hold jobs in memory while the queue database is unreachable - #810
IslamElsayed wants to merge 1 commit into
Conversation
|
The one red job, |
With buffer_enqueues_on_database_error enabled, an enqueue that fails because the database can't be reached holds the job in a bounded per-process buffer, reports it as enqueued, and enqueues it from a background thread once the database is back, as discussed in rails#549.
b6e5ae6 to
4a6f228
Compare
|
Update: I re-triggered CI by pushing the same commit again (identical tree). The flaky |
Closes #549, following the design @rosa agreed to there.
If the queue database goes down,
perform_laterraises and the job is lost. This adds an opt-in buffer: withconfig.solid_queue.buffer_enqueues_on_database_error = true, an enqueue that fails because the database can't be reached holds the job in memory. The job is reported as enqueued and gets enqueued from a background thread once the database is back.Behaviour
ActiveRecord::ConnectionNotEstablished(which also covers pool timeouts andDatabaseConnectionError) orActiveRecord::ConnectionFailed. Anything else,NoDatabaseErrorincluded, raises as today.perform_laterandperform_all_laterboth, as agreed in the issue. A bulk enqueue runs in one transaction, so it's held as a whole.successfully_enqueued?is true, andprovider_job_idstaysniluntil the job actually lands.enqueue_buffer_sizejobs (default 1,000). Past that, enqueuing raisesSolidQueue::Job::EnqueueErroras it does now.Job.enqueue_all, waiting 1s at first and doubling up to 30s. It keeps queue, priority,scheduled_at, concurrency keys and batch membership, and stops once the buffer is empty. It's created with the existingAppExecutor#create_threadand wrapped in the app executor.ActiveSupport::ForkTracker). That way it can't enqueue its parent's held jobs a second time, and it gets its own flusher, since threads don't survive a fork.buffer_enqueue(held, or rejected because the buffer is full),flush_enqueue_bufferandlose_held_jobs.Changes
SolidQueue::EnqueueBuffer(new, ~120 lines), plus tworescues inJob.enqueue/Job.enqueue_all.buffer_enqueues_on_database_errorandenqueue_buffer_size.Tests
test/unit/enqueue_buffer_test.rbcovers holding onperform_laterand on bulk enqueues, flushing with queue, priority and schedule intact, staying held while the database is still down, raising as before when disabled or for other errors, the size limit, the background flusher running and stopping, the loss report at exit, and a forked child starting empty.log_subscriber_test.rbcovers the three new log lines. I also broke each of the key behaviours on purpose (opt-in check, error filter, limit, requeue on failed flush, reporting as enqueued, bulk path, flusher stop, fork reset), and a test failed every time. The full suite passes on SQLite (the model and unit suites three times in a row), and CI covers MySQL and PostgreSQL.One point for review: when a flush fails with something other than a connection error, the batch is reported through
on_thread_errorand dropped rather than retried forever. Happy to change that if you'd prefer otherwise.I use AI assistance when working on patches, and I check and test everything before submitting.