fix: ES index upgrade never completes when several connectors share an alias - EXO-88232 - #769
Conversation
…n alias - EXO-88232 When two indexing connectors declare the same index_alias (e.g. notes pages and note versions on notes_alias), one ReindexESType task was submitted per connector. The first finished task saw the other one still pending and did nothing, but its finally block removed the whole alias entry from indexUpgrading, so the second task hit a NullPointerException and the alias was never switched to the new index (nor the old index deleted). - Track completion per connector; switch the alias and delete the old index once, after the last connector of that alias finishes. - Release only the finished connector's slot on failure, drop the alias entry when nothing is pending anymore. - Send the ES-side _reindex once per alias/pipeline instead of once per connector. - Add setExecutors/isIndexUpgrading for tests, and two tests covering the shared-alias case (mutation-verified against the previous logic). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
|
Second commit: a further test run on notes showed the main-thread side of the same problem — The new index is now (re)created once per alias ( |
…n alias - EXO-88232 (#769) ## Problem While upgrading `notes_v1 → notes_v2` (Meeds-io/notes#1756), the upgrade failed with: ``` ERROR | An error occurred while upgrading index notes_v1 java.lang.NullPointerException: Cannot invoke "java.util.Set.size()" because the return value of "java.util.Map.get(Object)" is null at ...ElasticIndexingOperationProcessor$ReindexESType.run(ElasticIndexingOperationProcessor.java:792) ``` Notes declares **two** connectors on the same `index_alias` (`WikiPageIndexingServiceConnector` and `NoteVersionLanguageIndexingServiceConnector`). One `ReindexESType` task is submitted per connector. The first task to finish saw `size() > 1`, did nothing, then its `finally` removed the **whole alias entry** from `indexUpgrading` → the second task NPE'd. Net effect: the alias was never switched to the new index and the old index was never deleted, on every alias shared by ≥ 2 connectors. ## Fix - Track completion **per connector** (`markConnectorUpgraded`); switch the alias and delete the old index once, after the last connector of that alias finishes. - On failure, release only the finished connector's slot; drop the alias entry when nothing is pending anymore. - Send the ES-side `_reindex` **once per alias/pipeline** instead of once per connector (it was duplicated for notes). - Logs name the connector; the pipeline-failure WARN now carries the exception. - `setExecutors(ExecutorService)` / `isIndexUpgrading(alias)` added for tests. ## Tests Two new tests in `ElasticOperationProcessorTest` with two connectors on `notes_alias`, run on a same-thread executor: - alias switch / old-index delete / `_reindex` each happen exactly once, after both connectors; - after only the first connector finishes, the alias is untouched and the upgrade is still pending. **Mutation-verified**: with the previous `size() > 1` + whole-alias `remove` logic put back, both tests fail; with the fix, the full `commons-search` suite passes. ## Classification N2 — shared indexing engine (cross-domain seam), no ACL/schema/REST surface touched. Knowledge: none — behaviour fix in the indexing engine, no domain-doc claim changes. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5 <noreply@anthropic.com>



Problem
While upgrading
notes_v1 → notes_v2(Meeds-io/notes#1756), the upgrade failed with:Notes declares two connectors on the same
index_alias(WikiPageIndexingServiceConnectorandNoteVersionLanguageIndexingServiceConnector). OneReindexESTypetask is submitted per connector. The first task to finish sawsize() > 1, did nothing, then itsfinallyremoved the whole alias entry fromindexUpgrading→ the second task NPE'd. Net effect: the alias was never switched to the new index and the old index was never deleted, on every alias shared by ≥ 2 connectors.Fix
markConnectorUpgraded); switch the alias and delete the old index once, after the last connector of that alias finishes._reindexonce per alias/pipeline instead of once per connector (it was duplicated for notes).setExecutors(ExecutorService)/isIndexUpgrading(alias)added for tests.Tests
Two new tests in
ElasticOperationProcessorTestwith two connectors onnotes_alias, run on a same-thread executor:_reindexeach happen exactly once, after both connectors;Mutation-verified: with the previous
size() > 1+ whole-aliasremovelogic put back, both tests fail; with the fix, the fullcommons-searchsuite passes.Classification
N2 — shared indexing engine (cross-domain seam), no ACL/schema/REST surface touched.
Knowledge: none — behaviour fix in the indexing engine, no domain-doc claim changes.
🤖 Generated with Claude Code