Skip to content

Hold jobs in memory while the queue database is unreachable - #810

Open
IslamElsayed wants to merge 1 commit into
rails:mainfrom
IslamElsayed:buffer-enqueues-on-database-error
Open

IslamElsayed wants to merge 1 commit into
rails:mainfrom
IslamElsayed:buffer-enqueues-on-database-error

Conversation

@IslamElsayed

Copy link
Copy Markdown

Closes #549, following the design @rosa agreed to there.

If the queue database goes down, perform_later raises and the job is lost. This adds an opt-in buffer: with config.solid_queue.buffer_enqueues_on_database_error = true, an enqueue that fails because the database can't be reached holds the job in memory. The job is reported as enqueued and gets enqueued from a background thread once the database is back.

Behaviour

  • Opt-in, off by default. With it off, nothing changes: the same errors raise the same way.
  • Only an unreachable database, meaning ActiveRecord::ConnectionNotEstablished (which also covers pool timeouts and DatabaseConnectionError) or ActiveRecord::ConnectionFailed. Anything else, NoDatabaseError included, raises as today.
  • perform_later and perform_all_later both, as agreed in the issue. A bulk enqueue runs in one transaction, so it's held as a whole.
  • Held jobs count as enqueued: successfully_enqueued? is true, and provider_job_id stays nil until the job actually lands.
  • Bounded: each process holds at most enqueue_buffer_size jobs (default 1,000). Past that, enqueuing raises SolidQueue::Job::EnqueueError as it does now.
  • Flushing: a single background thread retries through Job.enqueue_all, waiting 1s at first and doubling up to 30s. It keeps queue, priority, scheduled_at, concurrency keys and batch membership, and stops once the buffer is empty. It's created with the existing AppExecutor#create_thread and wrapped in the app executor.
  • Forks: a forked child starts with an empty buffer (via ActiveSupport::ForkTracker). That way it can't enqueue its parent's held jobs a second time, and it gets its own flusher, since threads don't survive a fork.
  • Exit: a normal exit tries once more. If the database is still down, it logs how many jobs were lost. A crash loses held jobs, the same trade-off as Sidekiq Pro's reliable client. The README says so.
  • Instrumented and logged: buffer_enqueue (held, or rejected because the buffer is full), flush_enqueue_buffer and lose_held_jobs.

Changes

  • SolidQueue::EnqueueBuffer (new, ~120 lines), plus two rescues in Job.enqueue / Job.enqueue_all.
  • Two settings: buffer_enqueues_on_database_error and enqueue_buffer_size.
  • README: the settings, and a short section under "Errors when enqueuing".

Tests

test/unit/enqueue_buffer_test.rb covers holding on perform_later and on bulk enqueues, flushing with queue, priority and schedule intact, staying held while the database is still down, raising as before when disabled or for other errors, the size limit, the background flusher running and stopping, the loss report at exit, and a forked child starting empty. log_subscriber_test.rb covers the three new log lines. I also broke each of the key behaviours on purpose (opt-in check, error filter, limit, requeue on failed flush, reporting as enqueued, bulk path, flusher stop, fork reset), and a test failed every time. The full suite passes on SQLite (the model and unit suites three times in a row), and CI covers MySQL and PostgreSQL.

One point for review: when a flush fails with something other than a connection error, the batch is reported through on_thread_error and dropped rather than retried forever. Happy to change that if you'd prefer otherwise.

I use AI assistance when working on patches, and I check and test everything before submitting.

@IslamElsayed

Copy link
Copy Markdown
Author

The one red job, Tests (3.3, sqlite, rails_7_2), failed in ReadyExecutionTest#test_claim_jobs_using_both_exact_names_and_a_prefix (2 jobs claimed instead of 6), and fail-fast cancelled the rest of the rails_7_1/8_0 jobs. It doesn't reproduce: running the full suite locally on the same Gemfile, database and seed (--seed 12042) gives 374 runs, 0 failures. Sibling prefix-claim tests in the same file have failed intermittently on other branches too (e.g. test_claim_jobs_using_queue_prefixes in run 34713920145, test_queue_order_is_respected_when_using_prefixes in run 36441240366). Could someone re-run the failed jobs? I can't from a fork.

With buffer_enqueues_on_database_error enabled, an enqueue that fails
because the database can't be reached holds the job in a bounded
per-process buffer, reports it as enqueued, and enqueues it from a
background thread once the database is back, as discussed in rails#549.
@IslamElsayed
IslamElsayed force-pushed the buffer-enqueues-on-database-error branch from b6e5ae6 to 4a6f228 Compare September 29, 2026 12:57
@IslamElsayed

Copy link
Copy Markdown
Author

Update: I re-triggered CI by pushing the same commit again (identical tree). The flaky ReadyExecutionTest passed this time, and every rails_7_2, rails_8_1 and rails_main job is green. The rails_7_1 and rails_8_0 jobs were cancelled after hitting the 15-minute job limit, but that isn't specific to this branch: other recent PR runs show the same timeouts. #36128048815 on stop-processes-whose-heartbeats-block has exactly this pattern (7_1 and 8_0 time out, 7_2/8_1/main pass), and runs on enqueue-jobs-on-shards, remove-dangling-recurring-tasks and fix-superfluous-select-on-single-job-schedule time out on every version except main. Locally, the full suite passes on the default Gemfile (Rails 7.1).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Missing DB Connection Resiliency

1 participant