Home Assistant calls turn_on regardless of current state, and applying a
scene re-asserts every captured state, so asserting the alarm on over a
rinse-only AddWashSet_1 rewrote it to _7 -- silently widening the moments
the user picked, with no state change on this switch to point at it. The
write is now refused when the mask is already non-zero.
Distinct from the off-then-on path documented in _add_wash_bit_switch,
where the appliance has no subset left to keep.
Also satisfies the ty check on the new tests: resolve(),
for_device_by_model() and exists_fn are all optional, asserted the way
the other capability tests do.
A DV6800N (DA_WM_A51_20_COMMON) reports the same /st/dryercourse/vs/0
courseTable 'Table_00' as issue #357's DVE45R6300W/A3 -- also
DA_WM_A51_20_COMMON -- and its reporter listed 14 selectable courses
read straight off the appliance's own menu, in the same order
/course/vs/0's supportedOptions enumerates them.
Initially treated this as a second, incompatible Table_00 code family
and added a device-model-keyed disambiguation layer to
laundry.cycle_select. That was an unproven assumption: the two
reporters' confirmed sets are mostly non-overlapping subsets (11 vs 14
codes), which is exactly what you'd expect from two models on a shared
board exposing different subsets of one course table via their own
/course/vs/0 supportedOptions -- not evidence of two different code
dictionaries. The one code both reporters confirmed, 'a5', means
Bedding on both, which is corroborating, not neutral. No confirmed
code conflicts between the two sets, so this folds #394's 14 codes
straight into the existing dryer_cycle_table_00 catalog entry
(mirrored to all shipped languages), same shape as #357's original
translations-only change.
Adds a scrubbed fixture + golden for the DV6800N dump and tests
confirming zero unbound hrefs and that its course codes resolve
through the shared, now-larger dryer_cycle_table_00 catalog.
on_notification discards on the DTLS reader thread. Snapshotting
_notified then assigning fallback_hrefs after releasing the lock
let a notify in that window land on the set object being replaced,
so a just-pushed href was classified silent until its next notify.
The master alarm switch wrote AddWashSet_7/_0 without consulting
_add_wash_mask, so a device reporting a wider mask (AddWashSet_15) or a
non-numeric one had it truncated to three bits -- the write that mask's
own docstring rules out, while the per-moment switches already refused it.
Also corrects the washer_ww6500 golden docstring: the fixture carries no
/oic/d and the test passes no device_types, so the device is typed solely
by the WW consumer prefix in its description. A51 is not a board token,
which makes that the fragile route worth naming.
Setup already starts `_run_subpolls` as a background task. Patching
`_poll_hrefs_blocking` then captured its unfiltered hot/warm batches
alongside the silent-only ones the test asked for.
on_notification discards from fallback_hrefs so a late first push
(after the 80% quorum snapshot) self-corrects instead of staying on
the 3s GET cadence. Empty subpoll slots skip the session lock.
Hoist airconditioner._has_sensor_type to common.has_sensor_type next to
sensor_item_value, add the issue #127 stub-rep carve-out, and match the
AC family's enabled_default=False on the purifier CO2 entity.
A switched-off washer or dryer fails in the DTLS handshake, not in a
poll: `_poll_once` opens the session itself, so `_connect_session` runs
to its 12s timeout with nothing to show. The poll path treated that like
any other poll failure and ran its reconnect -- close the session, pause,
poll again -- but there is no session to close and no association for the
device to clean up, so the retry was the identical handshake five seconds
later. That cost 29s of every 30s interval, and the same again on every
setup attempt for an entry with no snapshot to load from.
`_poll_once` now records which of the two failed, and the poll path skips
the retry when the handshake is what never completed. A session that
opened and then broke still reconnects within the cycle.
The log was the half the reporters saw: an ERROR every cycle (plus a
WARNING once three "reconnects" piled up) for a state this integration is
built to sit through, which issue #269's reporter read as the integration
having failed. An outage now reports once, DEBUG for the cycles after it,
and INFO when the device answers again.
Fixes#269
A live dump showed DrumCleanLog_'s newest entry moving every 30-90s on
its own, including well after a cycle had already finished -- it never
settles on a value worth showing. drum_clean_cycles_remaining, read from
separate WashingTimes_/DrumCleanProposal_ counters, is unaffected and
stays wired.
The washer/dryer readers this was originally built from (issues #9,
#258) aren't touched -- no report of the same churn there.
The progress sensor's catalog had no entry for the "Sanitizing" stage
Samsung dishwashers report, so it fell back to the raw device code
instead of a translated label. Add "sanitizing" to the progress
state table in every shipped locale.
Issue #92: try_enter_observe_mode treated every subscribed href as
push-covered once SUCCESS_FRACTION cleared, including ones that never
notified. Those then skipped _run_subpolls for the rest of the session.
Record subscribed-but-silent hrefs on fallback_hrefs (idle in observe
mode) and keep just that set on the hot/warm cadence.
Issue #387's TP1X_DA-AC-AIR-class board reports a CO2 item on
/sensors/vs/0. Same field/shape air_monitor.SENSORS already models
(carbon_dioxide / ppm). Gated with exists_fn so boards that don't
list the type (every current fixture) don't grow an empty entity.
The dishwasher's /course/vs/0 options[] array carries the same
WashingTimes_/DrumCleanProposal_/DrumCleanLog_ trio already read by
washer.py (issue #9) and dryer.py (issue #258), confirmed against a
live dump, but dishwasher.py never wired the shared
laundry.drum_clean_cycles_remaining/drum_clean_last_cleaned readers
in. Add the two sensors to CYCLE_OPTIONS the same way, plus tests and
an updated golden fixture.
CONTRIBUTING.md asks for one or two sentences over an essay, a pointer
rather than a re-derivation, and module docstrings that orient rather than
document the design. The identity change was written before that guidance
was read, and its comments narrate the reasoning at roughly twice the
length the conclusions need.
Cut to the conclusion and the evidence that makes it credible, keeping
every issue reference: _resolve_identity's docstring 41 lines -> 22 (now
shorter than the two longest already in that file), rekey_entry 23 -> 14,
its module docstring 21 -> 10, resolve_device_key 24 -> 16,
is_usable_device_id 15 -> 8, and the same treatment for the inline blocks
in _resolve_identity, the CONF_DEVICE_KEY note, the redaction rationale,
and the migration suite's module docstring.
Comments only; the suite and the thirteen-mutation sweep are unchanged.
CI runs `ty check custom_components tests`; the identity work was only
checked against `custom_components`, so eighteen diagnostics in the new
tests went unnoticed until the PR run.
Two shapes, both in the new files. `dev_reg.async_get` and
`ent_reg.async_get` return `... | None`, so chaining an attribute off them
is both a type error and, on the failure it exists to catch, an
AttributeError instead of the assertion that would say what went wrong --
replaced with `_device_identifiers`/`_entity_unique_id` helpers that assert
the row survived and return the field. And `_identity`'s `**kwargs`
dict-merge erased the field types it was constructing; spelling the three
optional fields out is clearer anyway.
No behaviour change; the full suite and the thirteen-mutation sweep still
pass.
Two real findings from review of the identity change.
A pre-v4 entry keeps its serial-keyed unique_id until its first live poll
adopts the UUID, which can be a long while for an appliance that is off
since the entry loads from its snapshot meanwhile. The config flow's UUID
check couldn't see such an entry, so re-adding the same appliance in that
window was waved through as a second entry -- and the two would collide the
moment the older one re-keyed, with rekey_entry resolving the collision by
deleting the duplicate rows, taking the original's entity_ids, history and
automations with them. The flow now also aborts on the legacy key, matched
together with the host so issue #381's two same-serial units at different
addresses stay separable, and gated on CONF_DEVICE_KEY being absent so the
check disappears once the entry has migrated.
Separately, the "nothing stored yet" branch adopted the polled identity
unconditionally, silently dropping the "same IP, different appliance" guard
_run_discovery has had since issue #236 -- and dropping it for precisely
the users who have been running longest. _resolve_identity now computes the
key an entry currently carries once and applies one corroboration rule to
it, so a pre-v4 entry defends itself exactly as a migrated one does.
Adoption still needs no corroboration where there is no identity claim to
defend: an entry keyed on its address (issues #83/#189) or with no stored
serial at all.
Two tests had to change with it, both because their setup was unfaithful
rather than because the behaviour regressed: the two-unit test now has its
devices report the shared serial they actually report, and the offline test
now models a placeholder-serial board, which is where an entry's key and a
snapshot's serial can genuinely disagree.
Three new tests cover the changed behaviour, and the mutation sweep is
extended to thirteen breakages -- including the two guards added here and
the host match that keeps the duplicate check from undoing #381's fix.
Also from review: rename three stale _resolve_key doc references to
_resolve_identity, and stop rebinding new_unique_id in rekey_entry.
The move onto the OCF device UUID rewrites the identity of registry rows
that a user's automations, dashboards, history and areas all hang off, and
it does so on the first live poll rather than inside async_migrate_entry.
That makes it the riskiest migration in the codebase and the one least
covered by its own tests, so it gets a dedicated suite organised around
what must not break rather than around the functions involved.
tests/localthings/test_identity_migration.py (19 tests) covers: the v3
upgrade end to end; the full v1 -> v4 walk an oldest install takes; a
host-keyed placeholder-serial board adopting a real identity; issue #381's
actual two-unit install migrating to separate identities; every
customization carried on an entity row (rename, icon, area, hidden_by) and
the device's own area; a composite appliance's subdevice identifiers and
via_device links; scoping to one config entry's rows; the key-boundary
match; upgrading while the appliance is off, and the deferred adoption
completing when it returns; restarting after the upgrade being a no-op;
and the three guards that defend an identity once it has moved.
tests/test_rekey_statistics_end_to_end.py drives a real in-memory recorder
to prove the claim the in-place rewrite exists to make: statistics are
filed under entity_id, so preserving the row preserves the history --
including on the one path that deletes a row rather than rewriting it.
It lives outside tests/localthings/ for the same reason the existing
statistics end-to-end test does (that package's autouse fixture starts HA
before recorder_mock can claim its database).
Every guarantee was checked by mutation: ten separate breakages of the
production code -- dropping the unique_id update, the subdevice prefix,
the entry scoping, the key boundary, the entity rewrite, both snapshot
guards, the no-demote rule, the rejected-serial rule, and recreating the
row under a new entity_id -- each fail the suite.
Two gaps this found and closed:
- ConfigFlow.VERSION had to move to 4. Home Assistant only calls
async_migrate_entry for entries behind the flow's version, so the v3 ->
v4 step silently never ran. Pinned with the reasoning written down.
- An offline load could re-key the registry from a stale snapshot, and
freeze that answer into CONF_DEVICE_KEY so the real UUID could never be
adopted afterwards. The guards existed; nothing proved they held.
The identity-migration tests move out of test_migration.py, which stays
about the migrations that finish inside async_migrate_entry.
Two Samsung air purifiers of the same model report the identical, well
formed serialNum `BS7SP9AW400114A` (issue #381). Since the entry's
unique_id, the device registry identifiers and every entity unique_id
were all minted from that string, the second unit was refused as already
configured, and would have collided entity-for-entity even if it hadn't
been.
This is the third firmware family to ship an unusable serialNum, after
`Nothing(SVC)` (#83) and the flash-unset sentinel (#189), and the first
one no heuristic can catch: the value is well formed, it's just shared.
`is_placeholder_serial` was a dead end.
So identity moves onto /oic/d's `di`, falling back to /oic/p's `pi`, then
the serial, then the host. `di` is what the protocol already uses to
address the endpoint -- if it were wrong or shared, OCF discovery and the
DTLS association wouldn't work at all -- and it's device-scoped, where
`pi` is platform-scoped and would be shared by a board hosting several
logical devices. Both units in #381 report a distinct `di`. A board that
answers neither resource lands exactly where it did before, so no
existing hardware regresses.
The re-key can't happen in async_migrate_entry: the UUID is only readable
from the device, and an entry can load entirely from its snapshot while
the appliance is off (#295). So v3 -> v4 only records the legacy key, and
the coordinator adopts the UUID on the first live poll, rewriting the
entity registry, the device registry (including subdevice identifiers)
and the entry's unique_id together. Rewriting rather than recreating is
what lets a user keep entity_ids, names, areas, statistics and every
automation that references them.
Three rules keep that adoption from misfiring:
- A poll that reads no UUID never demotes a UUID-keyed entry back onto
its serial, so one failed reconnect doesn't re-key every entity.
- A changed UUID is followed only when the serial still corroborates it
(a factory reset may regenerate `di`) or when the entry was keyed on
its IP, which was never an identity to defend.
- When the identity is rejected as a different appliance, the serial
isn't adopted either -- otherwise the intruder would gain exactly the
corroboration needed to win the next poll.
Also stop redacting `di`/`pi` from diagnostics. They're randomly assigned
per-unit UUIDs, not account data, and blanking them is what made the
first #381 diagnostics download unable to answer the only question it was
requested to answer. The owner-set device name stays redacted.
Fixes#381
- _on_cloud_courses_changed cleared the canonical-view cache but never
pushed state: select.py's current_option reads coordinator.data,
which only moves on async_set_updated_data, so a change here (the
new toggle, or apply_cloud_courses naming a program -- which has
called this same method since before the toggle existed) sat stale
in the UI until an unrelated poll or observe happened to run next.
Now calls _push_cache_snapshot() too.
- _refresh_cloud_course_issue now runs from __init__.py's
options-update listener on every entry save, not just a
cloud-course-specific one. Saving an unrelated option
(CONF_BYPASS_REMOTE_CONTROL, say) before this device's first poll,
or while it's rehydrated offline, read /course/vs/0 as empty --
indistinguishable from "nothing pending" -- and would delete a
Repair a real poll had every reason to raise. Now a no-op on an
empty rep, leaving whatever issue state already exists untouched
until a real poll can judge it.
- The "cloud_courses" menu's off-state note was a raw English string
built in config_flow.py and substituted via description_placeholders
into all 7 locales' descriptions -- unlike the SmartThings screen
names quoted elsewhere (deliberately English everywhere; that's a
third-party app's own label, not ours), this one named LocalThings'
own "Offer downloaded cycles"/"Device settings" labels, which are
translated per locale and should have matched. Replaced with a
permanent, state-independent sentence translated in the catalog
itself, in all 7 locales, instead of conditional Python-built text.
Also restores a word an earlier edit dropped from
async_step_cloud_courses's docstring ("can complete confidently").
Tests: two new regression tests, each confirmed to fail against the
pre-fix code before being fixed -- one drives coordinator.data through
a toggle via a fixture already sitting on a one-time cloud override, so
current_option actually depends on the cloud store instead of falling
back to the raw course code; the other simulates a second coordinator
against the same entry with an empty resource cache (a not-yet-polled
restart) and confirms an existing Repair survives an unrelated option
save. Full suite (1585 tests), ruff, and `ty check custom_components
tests` all pass.
Issue #364: several reporters got the "downloaded cycles not set up"
Repair despite never meaning to use the feature -- one device appears
to auto-populate a slot from a SmartThings-provided example. The two
reporters who did complete setup successfully both hit the same root
cause for their earlier failures: the SmartThings app has two
similarly-named screens ("Cycle", which lists everything including
local courses, and "Download cycles", the one that actually matters
here), and nothing in our instructions said to use the second one
specifically.
Global disable (CONF_CLOUD_COURSES_ENABLED, entry.options, default
on):
- New coordinator.cloud_courses_enabled property, mirroring
CONF_LEARN_MODES' shape -- off stops the Repair and stops offering
already-named programs as cycles, without discarding anything
already learned or named.
- Deliberately does NOT stop _observe_cloud_courses' passive recording:
guided/manual setup depend on live observation to detect a newly
selected program at all, and leaving it running means turning the
option back on immediately surfaces anything set up in the meantime
instead of asking the user to redo it. Documented on the const and
on the property.
- New __init__.py options-update listener calls a new
coordinator._on_cloud_courses_changed(), which both clears the
canonical-view cache (memoized, so a stale view would otherwise keep
answering with pre-toggle state -- caught by two failing tests
before this) and refreshes the Repair. Nothing else needed this
because every other option is read live on its own next use; cloud
courses is the only one with standing Repair/cache state to refresh
immediately rather than on the next unrelated change.
- Toggle exposed in Device settings as "Offer downloaded cycles",
alongside prose explaining why some devices show the Repair
unprompted.
Instructions, in every shipped locale (en/de/es/it/cs/nl/ko) --
otherwise a locale missing the new/changed strings would silently show
English or the old text, the same gap issue #376 already tests for:
- Every guided-setup screen, the manual edit form, and the Repair
itself now say explicitly: open the SmartThings app (not the
appliance), and tap "Download cycles" specifically -- a separate row
from "Cycle", further down the screen -- not the general cycle
picker. Also states plainly that the appliance doesn't need to be
nearby or running the cycle, just powered on and connected.
SmartThings' own screen names are kept in English in every locale
(verified only in the English app via the reporter's screenshots;
translating them without evidence of what Samsung's own localized
app shows would be a guess this codebase's translations otherwise
avoid).
- The Repair's description now also points at the new toggle for
anyone who doesn't want the feature at all.
- The "cloud_courses" menu screen shows a note when the option is
currently off, since guided/manual setup still work in that state
but nothing named there will appear as a selectable cycle until it's
turned back on.
Tests: coordinator-level tests cover the option defaulting on,
suppressing a new Repair, clearing an already-open one, hiding/
restoring the cycle-select entry as the option flips (which caught the
canonical-cache bug above), and that passive observation keeps running
regardless of the option. Options-flow tests cover the new field's
default and that it persists. Translation catalog tests
(test_every_language_mirrors_the_english_catalog et al.) cover every
locale's topology and placeholders for the changed/added strings.
Full suite (1583 tests), ruff, and `ty check custom_components tests`
(CI's exact invocation) all pass.
HA logs a removal warning (2027.8) every time this name is accessed on
releases that carry UnitOfDensity, attributed straight to this
integration since it's a plain module-level import. UnitOfDensity is
the replacement, but hacs.json's floor (2025.1.0) predates it existing
at all -- pytest-homeassistant-custom-component 0.13.316, the newest
available, still has no UnitOfDensity either, so this can't be a
static import on either branch without breaking support for part of
the version range.
Resolved with a runtime getattr instead: reads UnitOfDensity off the
homeassistant.const module if present and uses its
MICROGRAMS_PER_CUBIC_METER member, otherwise falls back to the plain
(un-deprecated, on those older releases) constant. The getattr
short-circuits before the deprecated name is ever touched on a release
new enough to have UnitOfDensity, so the warning stops firing there
without dropping support for anything still on the old one. Same
feature-detection shape _relabel_particulate_statistics already uses a
few lines down for new_unit_class.
Verified the resolution logic directly: against the installed HA
(2026.2.3, pre-UnitOfDensity) it resolves to the plain constant with no
warning; a simulated future homeassistant.const with UnitOfDensity
present resolves to it without ever touching the deprecated name (a
guard that raises on that access never fires).
Full suite (1573 tests), ruff, and `ty check custom_components tests`
(CI's exact invocation) all pass.
06's Korean text ('이불') is identical to the confirmed Bedding codes
24/6f, and Bedding reads better than the guessed 'XXL Laundry' wording
issue #342 originally gave it. Applied across all 7 locale catalogs
(matching each locale's own already-translated Bedding text, not a
fresh translation) and folded into issue #376's WF21T6500KV test as a
21st confirmed code instead of a flagged exclusion.
test_confirmed_washer_table_02_missing_course_names (#342) updated to
match; its docstring now notes 06's wording was later corrected by
#376 rather than pinning the old value as if still current.
Issue #376 reported Korean UI labels for 21 washer (Table_02) and 18
dryer (Table_03) codes that had no translation and were rendering as
raw hex in the UI, from a WF21T6500KV washer (DA_WM_A51_20_COMMON) and
DV19T8745BV dryer (DA_WM_TP1_21_COMMON).
Cross-checked every reported code against translations/ko.json before
translating anything: several share their exact Korean text with a
code the catalog already has a confirmed label for (washer '19'/'AI
맞춤세탁' matches '2b'/'69'; dryer '3a'/'살균건조' matches '21'; dryer
'3c'/'피트니스' even matches washer '2f', a cross-table reuse; etc.) --
those reuse the established label instead of a fresh translation. The
rest (Wool/Lingerie, Boil Wash, Soft Bubble, Padding Care, and others
with no catalog precedent) are new translations of the reporter's
Korean text.
One code is deliberately NOT applied: washer '06' ('이불', Bedding per
this report) conflicts with 'XXL Laundry', already locked in by
test_confirmed_washer_table_02_missing_course_names (issue #342). Two
reports of the same nominal Table_02 disagreeing on one code is a real
discrepancy, not a wording question -- left alone pending the reporter
(or another Table_02 owner) confirming which device's '06' is actually
wrong, same caution as the existing '24'/'33' transposition history
(issue #343).
Added to all 7 locale catalogs (en/de/es/it/cs/nl/ko), not just
English: HA falls back to English for any key a locale is missing, so
translations/en.json alone would still pass
test_every_language_mirrors_the_english_catalog's topology check while
leaving every other locale showing English text for these codes.
Tests: two new tests lock in the English labels and, for every reused
code, that every locale's label actually matches its anchor code (not
just English) -- the same gap issue #343 fell through, since the
topology test alone can't catch a locale-specific mistranslation.
Full suite (1575 tests), ruff, and ty all pass.
Two releases landed since the 0.1.6 pin, both confirmed by upstream
(QuiteYellow, in issue #361) as additive/opt-in with no interface
changes on our side:
- 0.1.7: server-certificate profiles (SamsungServerProfile), a bounded
DTLS handshake deadline (connect() now defaults to a 12s bound
instead of none), and a cancellable connect() via
ConnectCancellation. Our connect() call sites in coordinator.py and
config_flow.py pass no args, so they pick up the bounded handshake
for free; the cert-profile and cancellation pieces are opt-in and
unused here.
- 0.1.8: fixes blockwise OBSERVE notification reassembly
(QuiteYellow/SmartThings-Local#39) -- a notification carrying only
the first Block2 block was previously handed straight to
on_notification instead of being reassembled, and separately, the
Block2 loop could append a retransmitted/late block as if it were
the next one, or miscompute the next block offset after a mid-
transfer size downshift. Both corrupt a multi-block observed
resource without necessarily truncating it -- the "premature end of
stream" / "error decoding unicode string" CBOR failures reported in
issue #361 on /mode/vs/0. All error types stay within the existing
compatible-built-in table (ConnectionError/TimeoutError subclasses),
so no exception handling changes.
`>=0.1.6` already permitted pip to resolve 0.1.8 on a fresh install,
but an environment that already has 0.1.6 or 0.1.7 satisfying that
floor won't be upgraded by Home Assistant's requirement check -- which
is what #361's reporter is very likely still hitting on 0.22.0.
Raising the floor to >=0.1.8 forces that upgrade on the next release.
Verified against smartthings-local 0.1.8 from PyPI: full suite (1573
tests), ruff, and ty all pass. No source changes needed beyond the
three version pins (manifest.json, requirements-dev.txt, Dockerfile).
The OutdoorTemp_ options token was only surfaced on legacy boards
(is_legacy_board), even though 14 of 17 fixtures carrying the token are
non-legacy. issue #367 confirmed with a 48h field capture (r=0.92
against weather.forecast_home) that the token tracks real outdoor
temperature independent of board generation, and that no non-legacy
board exposes an alternative outdoor-temperature resource.
Split a token-presence-only exists_fn (_has_option_token_any_board) for
this token, leaving _has_option_token's legacy gate untouched for the
other options[] settings that still need it. The -55 offset itself was
only field-validated on Celsius-locale boards, so a second gate
(_reports_celsius, reading the board's own /temperatures/vs/0) keeps
the sensor off the one Fahrenheit-locale fixture on record rather than
guess whether the same offset and unit still apply there. Ships
enabled_default=False since multi-split installs report the same token
on every indoor head, which would otherwise create one duplicate active
sensor per head.
Updates the golden fixtures for the 13 affected Celsius-locale boards
and the artik051_krac test that had asserted outdoor_temperature stays
off newer boards; adds coverage for the Fahrenheit-locale gate.
The OutdoorTemp_ options token was only surfaced on legacy boards
(is_legacy_board), even though 14 of 17 fixtures carrying the token are
non-legacy. issue #367 confirmed with a 48h field capture (r=0.92
against weather.forecast_home) that the token tracks real outdoor
temperature independent of board generation, and that no non-legacy
board exposes an alternative outdoor-temperature resource.
Split a token-presence-only exists_fn (_has_option_token_any_board) for
this token, leaving _has_option_token's legacy gate untouched for the
other options[] settings that still need it. Ships enabled_default=False
since multi-split installs report the same token on every indoor head,
which would otherwise create one duplicate active sensor per head.
Updates the golden fixtures for the 14 affected boards and the
artik051_krac test that had asserted outdoor_temperature stays off
newer boards.
Widen async_rehydrate's guard to cover the identity and Subdevice rebuild,
not just the replay. A stored row missing a field the current dataclass
declares raised KeyError straight out of async_setup_entry, which only
handles ConfigEntryNotReady -- so the entry landed in SETUP_ERROR, which HA
never retries, with its DTLS session left open on the fixed source port the
next attempt binds. It now fails the same way an unreachable device does.
Write the snapshot immediately instead of through async_delay_save. A
deferred write outlives whatever queued it: removing an entry inside the
delay window deleted the file and then had it recreated, orphaned, when the
timer fired; and a reload scheduled by _reconcile_rehydrated read the
pre-reload snapshot back off disk, so a device going quiet again mid-reload
rehydrated the stale set and reconciled a second time. Banking it before the
reconcile fixes the ordering. Failures are logged rather than raised -- a
board reporting something the JSON encoder rejects must not break polling.
A coverage gap is a claim about what the device currently reports, so
replaying a discovery snapshot shouldn't make it. Offline it would restate
the last live poll's conclusion while pointing the user at a diagnostics
download that stays empty until the appliance answers, and any drift in the
resolved device name between snapshot and live would churn the issue.
Not deduplication: HA already keys issues on (domain, issue_id), preserves
dismissed_version across async_get_or_create, and reloads non-persistent
issues with their dismissal intact -- one row per entry, and an "Ignore"
survives restarts.
An appliance switched off at the wall used to take its whole config entry
down with it: async_setup_entry raised ConfigEntryNotReady, so the device
read as failed and its entities existed only as registry rows until the
appliance came back.
Loading the entry anyway isn't enough on its own. Entities here are the
output of discovery, discovery only runs inside a successful poll, and
platforms enumerate `bound` exactly once at forward time -- so an entry
that loads while offline loads empty, and with no listeners subscribed the
base coordinator stops rescheduling and never polls again.
Bank the resources dict each successful first cycle hands _run_discovery,
along with the subdevice candidate list and the /oic identity that route
the registry, and replay it through _run_discovery when the first refresh
fails. Storing the poll input rather than a rendered entity list keeps one
implementation of discovery instead of two: the offline entity set is
produced by the same code that produced the live one.
Three things fall out of that:
- Platforms judge entity existence against `discovery_resources`, not the
live cache. The live cache deliberately stays empty, which is what keeps
a restored entity `unavailable` rather than rendering a stale value for
an appliance nobody can currently reach.
- A live discovery that disagrees with the snapshot reloads the entry --
platforms can't adopt a changed set in place, so a firmware update or a
sibling subdevice that starts answering needs a fresh setup.
- The entry holds one coordinator listener for its lifetime, so polling is
scheduled regardless of how many entities are live.
An entry that has never reached the device has no snapshot, keeps raising
ConfigEntryNotReady, and closes its session on the way out as before -- no
metadata to build a device from, and it leaves room for setup flows that
need to interact with the appliance (#168).
Restores the two tests PR #303 rewrote, narrowed to that no-snapshot path.
PR #303 loads the entry when the first poll fails. Measured on its
branch, that produces an entry with zero bound entities and zero
coordinator listeners, so DataUpdateCoordinator never reschedules and
the device never recovers without a manual reload.
Record why entities can't be created offline here (discovery is the only
source of `bound`, and platforms enumerate it once), what a working
version would need (persisted discovery snapshot, reconcile-on-reconnect,
a listener that keeps polling alive), and the cheaper retry-and-reload
option that solves the filed issue on its own.