Our VitalPBX system experienced call failures. A saved core dump confirms a circular lock wait inside app_queue.so. Restarting Asterisk restored processing.
Environment
-
Asterisk packages
20.21.0-1 -
shared_lastcall=yesduring the incident -
Both involved queues have zero wrap-up time
-
Local-channel members using extension-state hints
Evidence
Thread 5: T11_Q3340, member 3342
Holds queue 3340’s lock; waits for queue 3140’s lock.
Thread 7: T11_Q3140, member 3142
Holds queue 3140’s lock; waits for queue 3340’s lock.
Both symbol-enabled stacks show:
extension_state_cb — app_queue.c:3050
update_status — app_queue.c:2804
update_queue — app_queue.c:6280, locking qtmp
The deadlock blocked 64 taskprocessor workers. Processing counters stopped advancing while Stasis backlogs grew.
The apparent locking issue is that status callbacks retain their original queue locks while shared_lastcall processing locks other queues before checking member membership. This appears to allow unrelated queues to deadlock, even with zero wrap-up time.
Recovery and workaround
Restarting Asterisk restored service. I subsequently changed the following in a custom configuration file:
[general](+)
shared_lastcall=no
Runtime verification, persistence after Apply Changes, and recurrence prevention have not yet been confirmed.
Questions for developers
-
Is this a known locking issue, and is a supported fix or backport available?
-
Is disabling
shared_lastcallthe recommended workaround? -
Is this custom configuration the supported way to preserve the setting in VitalPBX?
I have retained the core, symbol-enabled backtraces, and mutex-owner report. I do not have a deterministic reproduction.