Repository navigation
PyList_SetItem is ~2.7x slower per item on the free-threaded build #158660
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or errorperformancePerformance or resource usagePerformance or resource usage
on Oct 3, 2026 I think we might be able to add a fast path to
PyList_SetItemthat checks_PyObject_IsUniquelyReferencedto see if the list is shared and if not, doesn't acquire the critical section. Of course that adds a small overhead for shared lists, but I would think unshared lists are a lot more common.Per @kumaraditya303 in a chat message with me:
Avoiding locks by checking for unique reference is a bad idea in general, in many cases there is a single ref to object but it gets mutated by many threads concurrently because of borrowed refs.
So that won't work.
- addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Oct 3, 2026 My 2 cents:
I think it's the downside of the Stable ABI: stability at the cost of performance and I'm afraid we can't really do anything here. I think that's the tradeoff users need to accpet when using the stable ABI (at least according to the docs, AFAIU).
So what leaves us is:
- improve the interpreter overall so that this specific case isn't slower (I don't know if it's possible)
- silly idea that I haven't checkde at all: create a tuple and then call
PySequence_List. Depending on the size of your desired list, you might be faster (I haven't checked). But it can also be dramatically slower.
I would however say that
PyList_FromArraymay be a good (new) candidate (there is alreadyPyTuple_FromArraybut it's not in the Stable ABI either).An alternative is to add functions that are less safe in the stable ABI but I'm not sure the C API wg wants that. Having
PyList_SET_ITEMin the stable ABI (but not as a macro) would solve your issue I guess but that means we lose the safety the ABI was promising =/An alternative is to add functions that are less safe in the stable ABI but I'm not sure the C API wg wants that. Having PyList_SET_ITEM in the stable ABI (but not as a macro) would solve your issue I guess but that means we lose the safety the ABI was promising =/
I don't think we should do that.
silly idea that I haven't checkde at all: create a tuple and then call PySequence_List. Depending on the size of your desired list, you might be faster (I haven't checked). But it can also be dramatically slower.
This ends up being slower on the GIL-enabled build but faster under free-threading. It's a little unsatisfying because each extension needs
#ifdef Py_GIL_DISABLEDto individually work around this but it'll certainly help tokenizers in practice.This ends up being slower on the GIL-enabled
Yeah that's expected. How faster are we talking about here?
improve the interpreter overall so that this specific case isn't slower (I don't know if it's possible)
There is one way: PyO3's bindings could acquire a critical section on the unshared list and then CPython could add a second early check for recursive critical section acquisition that currently lives here:
cpython/Python/critical_section.c
Lines 24 to 45 in 81c50e9
// As an optimisation for locking the same object recursively, skip // locking if the mutex is currently locked by the top-most critical // section. // If the top-most critical section is a two-mutex critical section, // then locking is skipped if either mutex is m. if (tstate->critical_section) { PyCriticalSection *prev = untag_critical_section(tstate->critical_section); if (prev->_cs_mutex == m) { c->_cs_mutex = NULL; c->_cs_prev = 0; return; } if (tstate->critical_section & _Py_CRITICAL_SECTION_TWO_MUTEXES) { PyCriticalSection2 *prev2 = (PyCriticalSection2 *) untag_critical_section(tstate->critical_section); if (prev2->_cs_mutex2 == m) { c->_cs_mutex = NULL; c->_cs_prev = 0; return; } } } See https://github.com/python/cpython/compare/main...ngoldbaum:cpython:recursive-cs-extra-check?expand=1.
This amortizes the critical section acquisition cost over the whole list.
This unfortunately makes every other critical section acquisition that doesn't hit the existing check a little slower. Probably not worth it?
I checked and if PyO3's safe bindings acquire the critical section around the unshared list it actually already helps a bit, it's just not as fast as it possibly could be if we move the recursive critical section check. See PyO3/pyo3#6478.
Yeah that's expected. How faster are we talking about here?
~40% slower on the GIL-enabled build but ~30% faster on the free-threaded build. Both compared with initializing a list directly using
PyList_SetItem. So it doesn't buy you all the performance back but it does make the comparison a little better. Building a unshared list is only ~50% slower rather than 170% slower.I would however say that PyList_FromArray may be a good (new) candidate (there is already PyTuple_FromArray but it's not in the Stable ABI either).
Adding functions which would process multiple items per call sounds like an efficient approach and it would be a good API.
@ngoldbaum: What API would you need? Create a list from an array of objects? Set multiple list items from an array?
What API would you need? Create a list from an array of objects? Set multiple list items from an array?
In this case, the former. Something like the PyBytes_Writer API but for lists would help here.
Reacted by Petr Viktorin and Donghee Na
Bug report
Extensions built for the stable ABI can't use
PyList_SET_ITEM, so they fill new lists withPyList_SetItem. On the free-threaded build each call costs about 2.7x what it does on the GIL-enabled build, whilePyTuple_SetItemcosts about the same on both builds. Converting native sequences (pointer arrays, vectors) to lists is a hot path for bindings, so this becomes a visible slowdown when a project switches from abi3 to abi3t wheels. For example, in tokenizers a getter that returns three lists runs about 2x slower (see huggingface/tokenizers#2488 (comment)).Reproducer
fill.c:bench.py:Build one copy for abi3 and one for abi3t, then run them:
Results
3.15.0rc2, x86-64 Linux, gcc -O2:
CPython versions tested on:
3.15
Operating systems tested on:
Linux