The command-retry fix (issue #294) added a self._close_session() call to
async_send_command that isn't synchronized against _attempt_observe_mode,
which reads self._session and subscribes to it without holding
_session_lock. A write's retry racing an in-flight subscribe attempt
could tear down the session mid-subscribe -- or worse, land the close
*after* the attempt's grace wait already succeeded, letting it commit
observe mode against a session that's already gone: mode claims "Push"
forever, with nothing left to notice the underlying socket is dead.
Split ObserveManager.try_enter_observe_mode into four pieces
(subscribe_hrefs / await_observe_notifies / enter_observe_mode /
abandon_observe_attempt), keeping try_enter_observe_mode as a thin
wrapper so its direct callers in test_observe.py are unaffected.
_attempt_observe_mode now holds _session_lock only for the subscribe
burst (each send is fire-and-forget, not a network round trip) and
re-checks self._session is sess under the lock right before committing
-- sess keeps the old session object alive, so identity can't be
recycled onto a new one, which is what makes the check sufficient
without a separate generation counter. The wait itself stays lock-free,
so a command write is never blocked behind it.
Two more bugs the same investigation turned up, fixed in the same pass
since they're direct consequences of the design above:
- async_send_command's own successful reconnect didn't downgrade observe
mode the way the poll path's reconnect already does, leaving the same
stale-commit problem reachable with zero concurrency at all -- just a
write's retry succeeding while mode was observe. Replaced the poll
path's local just_downgraded_from_observe with an instance flag both
reconnect sites set, so either one triggers an immediate resubscribe.
- _maybe_retry_observe_mode's 600s throttle gated solely on
last_mode_change_ts, which _set_mode only stamps on an actual
transition -- a device that never succeeds at observe mode leaves that
timestamp stuck at construction time, so the throttle opens once and
never closes again, re-attempting on every single poll cycle instead
of every 600s. Now gates on the more recent of that timestamp and a
new _last_observe_attempt_ts, stamped on every attempt regardless of
outcome.