/diagnosis/vs/0 was the dump's only unbound href, now covered via
dishwasher.DIAGNOSIS (same shape already reused by airconditioner.py).
This steam-oven-class board's /mode/vs/0 options[] carries no UpperLamp_
token at all, unlike the NV7000BS-class board LAMP was proven against --
LAMP had no exists_fn, so it registered anyway, always read Off, and any
write to it was a no-op the device had no reason to honor. Gives it the
same options-token exists_fn gate issue #183 already added to
fast_preheat/natural_steam/energy_saving/cooktop_on_alert.
/alarms/vs/0 and /kidslock/vs/0 were the dump's two unbound hrefs -- both
are the exact shapes common.UNIVERSAL already models elsewhere
(common.ALARMS, common.KIDS_LOCK_VS_FALLBACK), picked individually rather
than pulling in all of UNIVERSAL to match this registry's existing
hand-picked-common style.
The six-vs-three burner count the reporter originally asked about is
expected behavior (the board's own /mode/vs/0 options genuinely advertise
six OperationState slots on hardware with three physical burners, with no
per-device signal to tell real slots from phantom ones) -- already
explained on the issue; this commit is scoped to the coverage warning.
/filter/airdustfilter/vs/0 was the dump's only unbound href -- this board's
internal deodorizing filter, same filterUsage/filterStatus field pair as
common.WATER_FILTER, but filterUsage here is already a 0-100 percentage
with no filterCapacity to divide by (confirmed by filterStatus=="wash" at
filterUsage=="100"). Uses air_-prefixed keys so a fridge with both a
water and an air filter gets two distinct entities.
This board routes purely via /oic/d's oic.d.airconditioner type (no
board-token match) and reports several resources the sibling
TP1X_DA-AC-CAC-01001 board (issue #191) left as a documented gap:
- /display/vs/0, /settings/sound/output/vs/0, /settings/sound/volume/vs/0
now reuse air_purifier.py's identical-shape capabilities instead of
duplicating them.
- /settings/sound/mode/vs/0 gets a new airconditioner.SOUND_MODE reading
the live supportedModes field, sharing laundry.py's existing
voice/tone/mute translation catalog since the value vocabulary matches.
- /csi/absenceclean/vs/0 and /csi/energysaving/vs/0 are new, genuinely
useful controls (absence auto-clean toggle, energy-saving mode select
plus its state/operatingStatus diagnostics).
- /dnd/autosleep/vs/0, /outdoorsharing/vs/0, /lifestyle/survey/vs/0,
/settings/sound/voice/vs/0 and /csi/information/vs/0 are ignored as
plumbing/unconfirmed data with no user-actionable state.
Also fixes a latent bug in air_purifier.SOUND_VOLUME: boards that report
minLevel/resolution but no maxLevel (this one) would have produced a
min=0/max=0 number entity instead of self-gating off.
Updates test_airconditioner_cac.py's documented coverage gap now that
sound_mode/sound_output/sound_volume are covered there too.
Review question on #304: a bare Comode_Off written over Comode_Sleep/Sleep_4
read back as Comode_Off/Sleep_0 at +8s and +38s, so the preset path cannot
leave a stale duration for the next nano selection to read as a running timer.
One map is not enough to judge by: with only OptionCode present, every
eoc-gated rule reads None, and None means the board does not publish the map
rather than that the feature is absent. artik051_dongle_fac_18k is exactly that
board and lost WindFree in every mode. Requiring both also keeps these bit
positions inside the family they were documented for -- the FAC and CAC dumps
carry only the older map, with values small enough that RAC positions read as
zeros.
Also from review: an unknown HVAC mode falls back the same way, Comfort is
spelled like the identical Speed rule, the unreachable AIComfort branch is
gone, the Cool code comes from the unit's own supportedModes, DlightCool gains
its catalog entry, and the Single User claim is dropped -- the app's own Single
User command sends Comode_Smart, so there is no distinct token to write.
coordinator is explicitly typed as LocalThingsCoordinator here, so ty
correctly sees _session's declared type (DtlsCoapSession | None) and
flags assigning a FakeObserveSession to it. The fixture's own
_connect_session replacement does the same swap without tripping ty,
but only because its self parameter is unannotated -- ty has nothing to
check the assignment against there. Deliberate here (this is the whole
point of the test: substitute a stand-in session), so silenced rather
than restructured; ty's --add-ignore confirmed the comment syntax
(ty: ignore[...], not the mypy-style type: ignore[...] used elsewhere
in this suite, which ty doesn't appear to honor for this rule).
A second review of the observe-mode phase split (previous commit) found
two real regressions it introduced, both in the same failure family it
was built to close:
- async_send_command's failed-retry branch closed the session, then
raised without downgrading observe mode -- the downgrade only ran on
the retry's success path. A retry that also fails still leaves the
session dead, so mode was left claiming "Push" on a session that no
longer exists, same as the bug this whole fix targets. Moved the
downgrade to run right after the close, unconditionally on how the
retry goes.
- _attempt_observe_mode's stale-session abandon (the identity-check
branch added in the previous commit) didn't flag a resubscribe. A
session swap discovered there means a fresh, never-tried session now
exists, but _last_observe_attempt_ts was already stamped for the
now-abandoned attempt -- so that new session sat unsubscribed for up
to _RECOVERY_RETRY_S (600s) instead of being retried on the next
cycle. Now sets _resubscribe_due, same as the two reconnect paths do.
Also closes two test-coverage gaps the same review surfaced by mutation
testing: no test asserted the lock actually holds during the subscribe
burst (only that it's released for the wait), and no test distinguished
the max() in _maybe_retry_observe_mode's throttle from using
_last_observe_attempt_ts alone -- both mutations left the full suite
green. Added one test for each, plus extended two existing tests for the
bug fixes above; all four confirmed via mutation testing (revert the
fix, watch the new/extended test fail; restore it, watch it pass).
One finding from the same review is intentionally left open: async_close
is the one self._session writer that doesn't take _session_lock, so a
close racing _attempt_observe_mode isn't covered by today's identity
check. This is pre-existing (the lock didn't cover any of
_attempt_observe_mode before this branch's earlier commits either), not
a regression from this branch's work, and is a shutdown/unload-path
question rather than the write-vs-observe-mode race this branch set out
to fix.
The command-retry fix (issue #294) added a self._close_session() call to
async_send_command that isn't synchronized against _attempt_observe_mode,
which reads self._session and subscribes to it without holding
_session_lock. A write's retry racing an in-flight subscribe attempt
could tear down the session mid-subscribe -- or worse, land the close
*after* the attempt's grace wait already succeeded, letting it commit
observe mode against a session that's already gone: mode claims "Push"
forever, with nothing left to notice the underlying socket is dead.
Split ObserveManager.try_enter_observe_mode into four pieces
(subscribe_hrefs / await_observe_notifies / enter_observe_mode /
abandon_observe_attempt), keeping try_enter_observe_mode as a thin
wrapper so its direct callers in test_observe.py are unaffected.
_attempt_observe_mode now holds _session_lock only for the subscribe
burst (each send is fire-and-forget, not a network round trip) and
re-checks self._session is sess under the lock right before committing
-- sess keeps the old session object alive, so identity can't be
recycled onto a new one, which is what makes the check sufficient
without a separate generation counter. The wait itself stays lock-free,
so a command write is never blocked behind it.
Two more bugs the same investigation turned up, fixed in the same pass
since they're direct consequences of the design above:
- async_send_command's own successful reconnect didn't downgrade observe
mode the way the poll path's reconnect already does, leaving the same
stale-commit problem reachable with zero concurrency at all -- just a
write's retry succeeding while mode was observe. Replaced the poll
path's local just_downgraded_from_observe with an instance flag both
reconnect sites set, so either one triggers an immediate resubscribe.
- _maybe_retry_observe_mode's 600s throttle gated solely on
last_mode_change_ts, which _set_mode only stamps on an actual
transition -- a device that never succeeds at observe mode leaves that
timestamp stuck at construction time, so the throttle opens once and
never closes again, re-attempting on every single poll cycle instead
of every 600s. Now gates on the more recent of that timestamp and a
new _last_observe_attempt_ts, stamped on every attempt regardless of
outcome.
An Opus review of the three prior commits on this branch (PR #306)
turned up two real defects and a documentation/test gap, all fixed
here:
- async_send_command's retry (issue #294) armed the write-settle
window before the retry existed, so the reconnect pause plus a
second PUT could eat into the time meant for the confirming poll,
reviving the revert-then-reapply symptom the window was sized to
prevent (issue #9). Re-arm it after a successful retry lands.
- test_send_command_reconnects_and_retries_after_socket_closed relied
on the observe-session fixture's no-op _close_session, so
self._session never actually went None and _do_put's reconnect
guard was never exercised -- the test passed even with that guard
deleted. Now overrides _close_session/_connect_session to actually
drop and rebuild the session, and asserts the reconnect happened.
- async_send_command's docstring still said "Fire-and-forget", which
stopped being true the moment it started retrying and raising.
Also trimmed the three comment blocks the review flagged as
reproducing their commit messages verbatim, per CONTRIBUTING.md's
comment-style rules.
One review finding is not addressed here and needs a decision: the
new _close_session() call in the command-retry path isn't
synchronized against _attempt_observe_mode, which touches the session
without _session_lock. A write's reconnect can race an in-flight
observe-mode subscribe attempt and tear down the session it's using.
Fixing it properly means broadening lock scope around observe-mode
entry, which risks blocking a write behind an up to ~15s subscribe
grace period -- a tradeoff not made unilaterally here.
A second finding (dropping the old .strip()'s per-line whitespace
handling in _normalize_pem) did not reproduce against a real
certificate/key, only against the test suite's placeholder PEM body,
so it's left as-is.
async_send_command's _do_put caught any exception, logged it, and
returned -- no reconnect, no retry, no error the user could see. A
command landing on a session Samsung's firmware closed between polls
(the same 'known device behavior' _async_update_data already
reconnects around) was silently lost, with nothing to do about it but
a manual reload of the device (issue #294).
Mirror the poll path's own recovery: on failure, close the dead
session, pause, and retry the PUT once against a freshly reconnected
one. If that also fails, raise a HomeAssistantError instead of just
logging, so the user gets a visible error rather than a command that
quietly did nothing. The retry runs under the same session lock the
poll path uses, so a write landing mid-reconnect can't race a
concurrent poll cycle rebuilding the same session.
When a poll fails and the immediate reconnect retry fails too,
_async_update_data returned the last-known snapshot as a degraded
success (issue #254) without ever touching observe mode. That's fine
for the data itself, but the connection-mode sensor reads straight
from self._observe.mode, and only the *successful* reconnect branch
ever changed it -- so a device that drops off the network entirely
(air-gapped, powered off, Wi-Fi down) left that sensor reporting
"Push" forever, hours after the session was actually dead (issue
#287).
Downgrade to poll mode on the failure branch too, without attempting
an immediate resubscribe: the reconnect that would normally justify
one just proved there's no live session to subscribe on. Recovery
still happens on its own once the device is reachable again, via the
existing poll-mode retry timer (_maybe_retry_observe_mode).
A PEM pasted from a text editor can carry bytes cryptography's parser
refuses outright: a UTF-8 BOM some Windows editors silently prepend,
CRLF line endings, and a stray blank line a paste can introduce
between the header/body/footer. None of those are meaningful in PEM,
but any of them surfaces as an opaque InvalidHeader with no hint of
what's wrong -- which is why the same certificate pasted from
Command Prompt's `type` (no BOM, no stray blank lines) loads fine
while the same file opened in an editor and copied doesn't (issue
#291).
Normalize at the point the pasted blob is first captured, not just
before minting the leaf cert: the same string is stored in the config
entry and reused to re-mint the leaf on a future reconfigure, so a
raw copy would keep failing every time it's read back, not just on
the first attempt.
The fixed list of six was offered in every mode on every legacy board. The
appliance publishes what it has as two bit maps in /mode/vs/0's options, and its
own app gates each comfort mode on a bit plus the current mode; this transcribes
that logic. WindFree also needs the mode written before it in Auto, which is
measured rather than assumed.
Sleep_<n> written on its own is answered 2.04 Changed and then discarded, so
the Number wrote nothing at all. Nano wind shares the same Comode_ slot, which
is why writing the nano preset over a running timer silently changed its
duration, and why the two sleep codes the board reports had to become presets:
a preset_mode outside preset_modes is not a state HA allows.
The Sleep_ token was published as if its value were hours. It is not: the
appliance's own app pairs a duration picker with the values it puts on the wire,
one to one, and the pairing is half hours.
0:00 0:30 1:00 1:30 2:00 2:30 3:00 4:00 5:00 ... 12:00
0 1 2 3 4 5 6 8 10 ... 24
So the entity capped at 12 hours actually set six, every value asked for was
halved on the appliance, and twelve hours -- the app's own maximum, stated in its
help text -- could not be reached at all. The reading is halved and the write
doubled, and the step drops to 0.5 because that is the resolution the picker
offers.
Half-hour steps are what the app offers below three hours; above that it offers
whole hours only, so a half hour up there is untested rather than known-bad. A
Number cannot change step part-way, and turning this into a Select of the app's
sixteen values would change the entity's domain on every unit that already has
one, so the step stays 0.5 throughout and the comment says why.
The descriptor's own comment used to admit the upper bound was a guess ("only 0
has been observed on hardware"). The guess of 12 was right; the unit it was
expressed in was not.
test_every_language_mirrors_the_english_catalog is right to fail on this: a key
present only in English falls back silently at runtime, which reads as a
half-finished translation rather than a missing one.
Also fills in the same key for ko.json, which the original fix predates:
Korean's translation catalog was added after this branch was cut and never
picked up filter_time_reset either.
FilterCleanAlarm_Clear, through the same single-token options merge as every
other setting on /mode/vs/0. Measured on an ARTIK051_KRAC_18K: 2.04 Changed and
FilterTime_95 (9 h 30 min) -> FilterTime_0, still zero on a fresh DTLS session
and on every poll after; none of the other 17 tokens moved and the alarm
entries stayed Deleted.
The counter has had no reset until now, and the descriptor said so: two earlier
rounds against live hardware failed, and the conclusion drawn from them was
that the reset had to be cloud-only. That conclusion was wrong, and the way it
was reached is the interesting part -- it came from diffing every resource the
appliance reports before and after pressing reset in Samsung's app, which
showed only the counter zeroing and the alarm clearing. A trigger token cannot
show up in such a diff, because a trigger is never stored. The appliance's own
app sends this token and skips the write when the counter is already zero.
Both failures stay in the comment, because they say what this is not: writing
FilterTime_0 (the value is not writable -- 5595 -> 5595 after 69 s, 1925 ->
1925 after 65 s, two units, opposite power states), and POSTing the cloud
capability's command name to /actions/vs/0 (real name, wrong transport).
Gated on the FilterTime_ token, so it appears only where there is a counter to
reset; newer boards report filter usage through their own resource and would
need a different mechanism.
ko.json (PR #283) was written against an older en.json and never picked up
the ten keys two later PRs added: the AI Purify sensing entities
(switch.periodic_air_sensing, switch.periodic_sensing_skip_status,
number.sensing_interval, select.sensing_mode + its three states,
time.sensing_skip_start/end) and button.auto_clean_stop. Merging main
into this branch surfaced the gap via
test_every_language_mirrors_the_english_catalog.
Translated the missing entries, matching the terminology and phrasing
ko.json already uses for adjacent keys (e.g. periodic_air_sensing's
existing binary_sensor entry, auto_clean's "자동 청소"), and kept "AI
Purify" as the untranslated brand name the same way cs/es/it/nl do.
ko.json's topology now matches en.json's exactly; full suite (1213
tests), ruff, and ty all pass.
CONTRIBUTING.md gains a "Code comments" section: comment the why not
the what, keep it to a sentence or two with a pointer to the load-bearing
evidence, don't re-derive a sibling's already-documented reasoning, and
move failed-attempt investigation logs out of inline comments.
Applied that policy across the codebase: condensed sprawling module
docstrings, per-entity essays, and multi-paragraph rationale blocks down
to their load-bearing conclusions, while preserving the actual "why"
(issue numbers, calibration evidence, gotchas, don't-guess rationale).
No functional code changed — verified via diff review, ruff, ty, and the
full pytest suite (1211 passed).
One inline investigation log (the AC filter-reset "tried and failed"
notes) moved to docs/investigations/ac-filter-reset.md rather than being
deleted, per the new guideline on where that kind of record belongs.
Three tokens describe the drying cycle these boards run after cooling, and the
switch only covered the first. AutocleanProgress_ is how far a running cycle
has got, and StopAutoClean_ is a channel for ending one early -- its presence
is what says the appliance accepts that at all, which is how the appliance's
own app gates its stop button. Both tokens are reported by the ARTIK051_KRAC_18K
that issue #136 was about, and by every KRAC fixture here.
The percentage scale is the app's own: it renders the token into a
`<progress max="100">` with a "{{value}}%" label beside it. An idle unit reports
1 rather than 0 -- the same floor the laundry firmware's progressPercentage sits
at when Ready -- so 0-vs-1 is not a reliable "is it running" test, and the
button is deliberately not gated on it.
The sensor shares AUTO_CLEAN's catalog entry the way auto_clean_legacy already
shares the switch's: same figure, different board generation, distinct key so
nothing collides if a board ever reported both.
Stacked on #289 (this branch is cut from it) -- rebase or merge that first.
The frozenset was a parallel structure for a per-row fact, and the comment
above the tuple already described it as a fourth column -- so the comment
promised the right shape and the code did something else. Fixed to the shape
the comment described: _AIR_QUALITY_SENSORS carries state_class per row and
the comprehension unpacks it, with _RECORDED_AIR_QUALITY and its duplicated
rationale block deleted.
air_monitor imports the same rows and now unpacks four, but discards the
fourth. That board (issue #210) has stamped all five readings as
`measurement` since it was added; consuming the column would silently drop
long-term statistics for Odor and CleanLevel on shipped devices, which is a
behaviour change this branch has no evidence to make. The grade/concentration
split stays scoped to the air purifier.
test_shared_sensor_tuple_keeps_its_three_column_shape asserted the premise
this replaces -- that widening the tuple breaks air_monitor's import -- so it
is replaced rather than renumbered: one test that the rows carry their own
state_class, and one that air_monitor still imports and still stamps all five.
Dropping native_min to 0 fixed the read range and quietly opened a write:
native_min governs what the user can enter, not just what renders, so 0
became enterable and would have gone out as periodicSensingInterval "0".
Nothing establishes what that does to this board -- both fixtures report 600,
the app's smallest choice is 10 min, and 60 s is the lowest value confirmed
accepted. The two precedents leaned on differ in exactly the way that matters:
oven.cook_time and operational.delay_start_hours sit at a zero floor under a
value where 0 is a real setting ("no timer", "no delay").
One minute is also the resolution this board reports results at.
lastSensingTime lands on an exact minute on every sample from the AVT-WW-TP1
and A-VTWW-TP2 boards -- both fixtures, plus eleven consecutive live readings
-- where the TP1X/AC/hood boards report arbitrary seconds. A sub-minute
interval is unobservable here whether or not the board honours it.
So native_min goes to 1 rather than 0, and the read rounds up instead of to
nearest so a sub-minute reading renders as 1 rather than falling below the
entity's own floor. The write still refuses anything under a minute -- a None
return, the silent no-op range_hood._lamp_level_write uses for a level the
device didn't advertise -- since native_min only guards the UI path, not a
service call.
Review feedback on #268. The largest change is that the sensing-mode select
no longer folds two device fields into one control.
periodicSensingActivationState and autoExeState are independent knobs, and the
appliance presents them that way -- its own UI has an on/off for AI Purify
separately from the three mode choices. Folding them lost two things: a
configured action was invisible while the feature was off, and no select
option could toggle the feature without also overwriting the action. The
switch was not the duplicate it looked like.
So the switch now owns periodicSensingActivationState alone, and the select
owns autoExeState alone. That resolves the hardcoded-options finding at the
source rather than working around it: the select reads supportedAutoExeState
via options_field -- the same shape SOUND_MODE already uses for
supportedModes -- instead of carrying a typed-in tuple, so a board advertising
a fourth action is accepted on both the options list and the write path.
_sensing_mode, _sensing_mode_write and _SENSING_MODE_BODIES are all gone with
the fold.
Option slugs are now the advertised values lowercased (off / airpurify /
alarm) rather than invented names. The catalog carries the labels, so the two
'off's stay distinguishable in the UI: the switch's means the unit isn't
sampling, the select's means it samples and doesn't act on the reading -- what
the app calls "sensing only".
Also from the review:
* _interval_minutes checks `is None` so a reported 0 stays 0, and native_min
drops to 0 since sub-30s values round there. oven.cook_time and
operational's delay hours are the precedent -- both convert a device time
value and floor at zero. The Number-rather-than-Select choice is now
stated in the write helper: the app offers three fixed intervals, but this
resource advertises no supported-values or range field (supportedAutoExeState
sits right beside it, so the board does advertise constraints where it has
them) and it accepted 60 s, six times finer than the app's smallest choice.
* _skip_time_write no longer splices a malformed half back onto the wire.
The read side already refuses one it can't parse; the write side now
zeroes it to match.
* air_sensing_state and last_air_sensing_level lose enabled_default=False,
matching range_hood.AIR_LEVEL_CHECK. Hiding two of three read-only keys
while claiming key parity with that capability -- and leaving the third
visible -- had no justification behind it.
* The catalog-parity test drops periodic_air_sensing from its key set: that
key is a SwitchDesc here and a BinarySensorDesc on the hood, so the two
live in different platform catalogs and are worded differently. The claim
now covers only the three read-only sensor keys, where it holds.
* Tests route through the descriptors (_desc(key).write_fn / .value_fn)
rather than module-private helpers, matching test_air_monitor_capabilities.
startSensingOnce stays unbound, now explicitly rather than by omission -- the
module comment records it as deferred. It looks like a one-shot "sense now"
button, but this board acknowledges writes it discards, and nothing has
confirmed the side effect yet.
Goldens are untouched: the key set is unchanged, only sensing_mode's value
moves from the folded slug to the raw autoExeState.
/airlevelcheck/vs/0 has been covered as "periodic air-quality sensing
scheduler plumbing" since the registry gained a coverage stub for it. Two
AVT-WW-TP1-23-AXX500 dumps (issues #84 and #190) show it is not plumbing: it
drives the feature the SmartThings app calls AI Purify, where the unit wakes
on a timer, samples the air, and optionally acts on the result. Every field
is named, none are opaque, and two of them are already user-set on the
reported units.
The select's three on-states are the app's own options rather than an
invented grouping -- it offers exactly "Sensing only" (sample, take no
action), "Auto clean" (purify while the air reads bad, stop once it
improves) and "Get notified" (raise a SmartThings notification). Labels were
transcribed from the Korean app and rendered in English; the auto-stop half
of "Auto clean" is the app's own description and is not otherwise visible in
the dump, which reports only the selected autoExeState. The remaining entity
names follow their raw fields rather than inventing a concept -- the skip
window is "sensing skip", after periodicSensingSkipStatus/Time.
Three of this registry's four board families report the resource with the
same field names -- TP1X_DA-AC-AIR (#130), A-VTWW-TP2 (#151) and AVT-WW-TP1
(#84, #190). Only ARTIK051_TVTL (#56) has no such href, and its golden is
unchanged. Bound unconditionally rather than behind a match_fn; the one field
that genuinely varies (periodicSensingInterval, absent on the #130 board) is
gated per-entity, so that board gets eight entities instead of nine rather
than a broken one.
range_hood.AIR_LEVEL_CHECK already models this same href, and its read-only
keys are reused verbatim here so both families share one catalog entry. It is
deliberately not imported: the hood exposes periodic_air_sensing as a
read-only BinarySensorDesc and this board needs a writable SwitchDesc on that
key, so reusing the hood's capability would migrate every hood user's entity
to a different platform.
Every write was exercised on AVT-WW-TP1-23-AXX500 hardware. This board hands
out 2.04 for writes it silently discards (see HEPA_FILTER's filter-reset
note), so an echo proves nothing -- each was judged by whether the value
survived a reconnect, which forces a new DTLS session, fresh discovery and a
fresh observe of the href, leaving no cached state to read back:
* sensing_mode's combined two-field PUT lands both fields, both ways:
sensing_only -> auto_purify raises autoExeState with activation still On,
and back again lowers it.
* The sensing-skip switch holds Off -> On and back.
* The half-preserving time writes hold: from 13:00-23:00, writing
start=07:30 then end=22:00 left the device on '07302200' -- each write
kept the half it wasn't given.
* periodic_air_sensing and sensing_interval: writing 60 s drove an observed
~60 s sensing cycle.
* The read side of the skip window is separately cross-confirmed on two
units: #84's sits at the inert '00000000', #190's carries a real
'03002300' (03:00-23:00), which is what pins the HHMMHHMM split.
* The other two families get the writes on field-shape grounds -- the same
basis on which they already share MODE, HEPA_FILTER and the air-quality
sensors.
range_hood._timestamp moves to common.epoch_to_utc so both callers share it,
matching how filter_usage_percent was shared. No behaviour change.
Every existing entity is untouched: the three golden updates are purely
additive, no renames, no unit or device_class changes.
_serial_from_unique_id took the entry's unique_id at face value. That is
right for an entry whose unique_id holds a real serial, but the unique_id
records what the config flow believed when it ran, not what the registry
holds now -- and for two firmware families those are different things.
Entries added before the placeholder rules landed (issues #83/#189) were
keyed on the placeholder itself: `localthings_Nothing(SVC)` for the
ARTIK051_DONGLE_REF dongles, `localthings_FFFFFFFFFFFFFFF` for the
DA_WM_A51_20_COMMON laundry boards. The coordinator has been resolving
those same boards to the host ever since, so their devices and entities
are host-keyed today. Migration read the placeholder back off the
unique_id, decided the host-keyed rows were the stale ones, and rewrote
them onto the placeholder -- reintroducing exactly the collision those
issues exist to prevent, since every unit of the family reports the same
placeholder and would go back to sharing entity unique_ids.
Run the recovered string through resolve_serial, which is the whole point
of that helper being shared. The old `host:port` special case stays: it's
a config-flow-history artifact rather than a device-reported serial, so
resolve_serial can't recognize it.
The repair pass had a second, narrower way to lose data. Removing a device
takes its entities with it (entity_registry.async_device_modified), and
the removal branch ran after the entity pass -- so an entity that had just
been re-keyed rather than removed, because its serial-keyed key was free,
was destroyed a few lines later along with the entity_id, name and area
the rewrite existed to preserve. Move surviving entities onto the device
they now belong to before removing the duplicate.
Reachable when the serial-keyed device exists but a given entity's
serial-keyed key doesn't -- e.g. the user deleted the visible duplicate by
hand, which is the first thing anyone hitting #236 tries.
Also fold the modelNum `<model>|<board>` split into resolve_model beside
resolve_serial. The config flow and _run_discovery each had their own copy
under a comment promising they matched; a device that renames itself on
the first poll is what a drift there looks like.
The Czech catalog was the only file in the repo still using CRLF, which
made every edit to it show up as a whole-file rewrite in diffs and hid the
one line that actually changed.
Content is byte-identical apart from the line endings, and the file now
matches the exact json.dumps(indent=2, ensure_ascii=False) formatting the
other four catalogs already use.
Adding a device had one message for nearly every way it could fail: "Cannot
connect to the device. Verify the IP address is reachable and the CA
credentials are correct." That covers an IP with nothing on it, an
appliance on cloud-only firmware, a device still holding the session from
the last attempt, a device that answered and rejected our certificate, and
Home Assistant having no internet to reach Samsung's cloud. Only one of
those is fixed by checking the IP and the CA credentials, and the message
gave no way to tell which one you had.
The probe already gathers enough to tell them apart, so classify it:
- cert_rejected the appliance sent a certificate alert. The CA
credentials aren't the AC14K_M CA it trusts, or they
don't pair. Far and away the most common real setup
mistake, and previously indistinguishable from a typo
in the IP address.
- handshake_failed a fatal alert unrelated to the certificate (protocol or
cipher mismatch) -- no amount of fiddling with CA
credentials will fix it.
- handshake_timeout the ClientHello probe proved a DTLS server is right
there, but the handshake never finished. Usually the
appliance is still holding the association from a
previous attempt; it clears on its own in about a
minute.
- ports_closed ICMP port-unreachable on the whole range: something is
at that address and it isn't exposing a local API.
Cloud-only firmware (TCP 8888 only) lands here.
- no_dtls_server some ports open|filtered, none speaking DTLS -- likely
another device on that IP.
- no_response nothing came back at all.
- cloud_unreachable couldn't reach Samsung's cloud gateway for the UUID.
An internet problem on HA's side, not the appliance's.
- unexpected_response authenticated fine, then returned something we can't
read. Neither connectivity nor credentials.
Certificate alerts are read back out of the error text OpenSSL puts in
DtlsCoapSession's ConnectionError, not by re-probing. The library's
diagnostic probe would report the alert authoritatively, but it drives the
handshake far enough to commit association state on the device -- and an
orphaned association is exactly what makes the *next* attempt time out
(RFC 6347 4.2.8), which is a bad trade on a path the user is about to
retry.
Telling ports_closed from no_response needs the UDP sweep to separate a
refusal from an unreachable. Both leave a port "not live", but ECONNREFUSED
is a *response* -- the host is there -- while EHOSTUNREACH/ENETUNREACH mean
the datagram never left. A wrong IP on the local subnet never answers ARP
and fails every send that way, so treating the two alike would have told
those users their appliance was on cloud-only firmware. The sweep now
returns live/refused/unreachable separately, and the preferred-port rescue
moved out of it into _sweep_ports: the rescue is a candidate-selection
decision, and folding it into the sweep's verdict destroyed the evidence
the message is built from.
Every failure carries the error key that fits it, so the flow maps
exceptions instead of guessing, and logs the specifics (alert name, per-port
outcome, response code) at warning level -- the messages that mention the
log now have something to point at.
Certificate re-minting for a reused leaf is also narrower and more correct
as a result: it now triggers on CertRejected specifically, rather than on
"every attempt raised ConnectionError and a port was confirmed".
All five translation catalogs carry the eight new messages. The non-English
ones are my own work rather than a native speaker's; corrections welcome.
Two problems, one setup path.
Port detection (issue #211): the config flow found the DTLS port by
elimination -- a 1-byte UDP probe can't tell a silent port from a real
DTLS server, so every port it couldn't rule out got a full certificate
handshake, and every false positive cost the whole 12s HANDSHAKE_TIMEOUT_S
before the next was tried. Adding an appliance took 30-40s.
smartthings-local 0.1.2 ships a stateless ClientHello probe that settles
this positively: a real DTLS server answers with a HelloVerifyRequest in
~1 RTT, and per RFC 6347 4.2.1 it does so without allocating association
state, so the probe leaves nothing behind on the appliance. The whole
49152-49160 range is probed at once and exactly one confirmed port is
given a certificate handshake. Fanning out is safe here in a way racing
real handshakes is not -- each probe is bounded by a 3s budget, so the
pool costs one probe's wall clock rather than the sum of the range, with
no losing threads left running behind us.
The UDP sweep stays as the fallback for when the probe confirms nothing:
it errs in the opposite direction (it reports everything it can't rule
out), so it still surfaces a device on a path that eats our ClientHello,
and it keeps its issue #192 preferred-port rescue.
Port detection now runs first and needs no credentials, so an unreachable
host fails before any round trip to Samsung's cloud. And a second
appliance reuses the existing entry's leaf cert rather than re-minting --
every device accepts the same one -- which makes adding one independent
of Samsung-cloud reachability. A confirmed-live device rejecting the
reused leaf (the UUID does rotate) re-mints and retries once, so reuse
stays self-correcting; a timeout doesn't, since a fresh cert can't fix
nothing answering.
Identity (issue #236): the coordinator seeded device_serial with the
configured host and only replaced it after the first successful poll. But
device_serial mints *permanent* registry keys -- entity unique_ids and
device identifiers -- so anything registering before that poll returned
was written into the registry keyed on the IP address forever. The
connection-mode sensor is added unconditionally rather than from `bound`,
so it was the reliable victim: when the serial-keyed identity appeared
moments later HA created a second device and entity, and the IP-keyed
pair was orphaned. Deleting them didn't help; the next restart that lost
the race recreated them.
The probe already learns the identity, so store it on the config entry --
serial, model, manufacturer, device type. The coordinator seeds
device_serial and its DeviceInfo from those at construction, so keys are
correct from the first entity that registers even if the first poll is
slow or fails outright. There is no placeholder left to correct.
Discovery now treats the registered identity as authoritative rather than
re-keying a device that already has registry entries; it adopts and
persists the polled identity only for an entry that has none, and warns
if a different appliance answers on the same IP.
Entry version 1 -> 2 recovers the serial from the entry's unique_id (the
flow has always keyed it on the probe's serial) and repairs what the old
registration orphaned: IP-keyed devices and entities are rewritten in
place where the serial-keyed key is free -- keeping entity_id, name, area
and every automation referencing them -- and removed where both exist,
since the IP-keyed one has been dead since the restart that made it.
Placeholder-serial boards (issues #83/#189) were keyed two ways at once,
`host:port` on the entry and `host` in the registry; migration collapses
the entry onto the registry's form. One resolve_serial() now serves both
sides, so they can't drift apart again.
The remaining step in the desired pipeline -- probe for subdevices, then
register devices, then populate entities -- already holds:
_enumerate_subdevices_blocking runs before _run_discovery, which runs
before platforms are forwarded. Duplicating it in the config flow would
mean re-running Pattern B's per-href fallback probe, which is the
opposite of what issue #211 is about.