Commit Graph
555 Commits
Author SHA1 Message Date
Marc Billow 789aaf9849 Merge pull request #306 from mbillow/claude/issue-triage-backoff-xkeddl
Fix reconnect/retry gaps found in issue triage (#291, #287, #294)
2026-08-05 22:09:52 -04:00
Marc Billow e3e7f4f43c test: suppress ty's invalid-assignment on the fake-session swap
coordinator is explicitly typed as LocalThingsCoordinator here, so ty
correctly sees _session's declared type (DtlsCoapSession | None) and
flags assigning a FakeObserveSession to it. The fixture's own
_connect_session replacement does the same swap without tripping ty,
but only because its self parameter is unannotated -- ty has nothing to
check the assignment against there. Deliberate here (this is the whole
point of the test: substitute a stand-in session), so silenced rather
than restructured; ty's --add-ignore confirmed the comment syntax
(ty: ignore[...], not the mypy-style type: ignore[...] used elsewhere
in this suite, which ty doesn't appear to honor for this rule).
2026-08-06 02:07:39 +00:00
Marc Billow a3cc918343 fix(coordinator): two gaps a follow-up Opus review found in the split
A second review of the observe-mode phase split (previous commit) found
two real regressions it introduced, both in the same failure family it
was built to close:

- async_send_command's failed-retry branch closed the session, then
  raised without downgrading observe mode -- the downgrade only ran on
  the retry's success path. A retry that also fails still leaves the
  session dead, so mode was left claiming "Push" on a session that no
  longer exists, same as the bug this whole fix targets. Moved the
  downgrade to run right after the close, unconditionally on how the
  retry goes.

- _attempt_observe_mode's stale-session abandon (the identity-check
  branch added in the previous commit) didn't flag a resubscribe. A
  session swap discovered there means a fresh, never-tried session now
  exists, but _last_observe_attempt_ts was already stamped for the
  now-abandoned attempt -- so that new session sat unsubscribed for up
  to _RECOVERY_RETRY_S (600s) instead of being retried on the next
  cycle. Now sets _resubscribe_due, same as the two reconnect paths do.

Also closes two test-coverage gaps the same review surfaced by mutation
testing: no test asserted the lock actually holds during the subscribe
burst (only that it's released for the wait), and no test distinguished
the max() in _maybe_retry_observe_mode's throttle from using
_last_observe_attempt_ts alone -- both mutations left the full suite
green. Added one test for each, plus extended two existing tests for the
bug fixes above; all four confirmed via mutation testing (revert the
fix, watch the new/extended test fail; restore it, watch it pass).

One finding from the same review is intentionally left open: async_close
is the one self._session writer that doesn't take _session_lock, so a
close racing _attempt_observe_mode isn't covered by today's identity
check. This is pre-existing (the lock didn't cover any of
_attempt_observe_mode before this branch's earlier commits either), not
a regression from this branch's work, and is a shutdown/unload-path
question rather than the write-vs-observe-mode race this branch set out
to fix.
2026-08-06 02:01:30 +00:00
Marc Billow 69f93be4dc fix(coordinator): close the observe-mode race an Opus design review found
The command-retry fix (issue #294) added a self._close_session() call to
async_send_command that isn't synchronized against _attempt_observe_mode,
which reads self._session and subscribes to it without holding
_session_lock. A write's retry racing an in-flight subscribe attempt
could tear down the session mid-subscribe -- or worse, land the close
*after* the attempt's grace wait already succeeded, letting it commit
observe mode against a session that's already gone: mode claims "Push"
forever, with nothing left to notice the underlying socket is dead.

Split ObserveManager.try_enter_observe_mode into four pieces
(subscribe_hrefs / await_observe_notifies / enter_observe_mode /
abandon_observe_attempt), keeping try_enter_observe_mode as a thin
wrapper so its direct callers in test_observe.py are unaffected.
_attempt_observe_mode now holds _session_lock only for the subscribe
burst (each send is fire-and-forget, not a network round trip) and
re-checks self._session is sess under the lock right before committing
-- sess keeps the old session object alive, so identity can't be
recycled onto a new one, which is what makes the check sufficient
without a separate generation counter. The wait itself stays lock-free,
so a command write is never blocked behind it.

Two more bugs the same investigation turned up, fixed in the same pass
since they're direct consequences of the design above:

- async_send_command's own successful reconnect didn't downgrade observe
  mode the way the poll path's reconnect already does, leaving the same
  stale-commit problem reachable with zero concurrency at all -- just a
  write's retry succeeding while mode was observe. Replaced the poll
  path's local just_downgraded_from_observe with an instance flag both
  reconnect sites set, so either one triggers an immediate resubscribe.

- _maybe_retry_observe_mode's 600s throttle gated solely on
  last_mode_change_ts, which _set_mode only stamps on an actual
  transition -- a device that never succeeds at observe mode leaves that
  timestamp stuck at construction time, so the throttle opens once and
  never closes again, re-attempting on every single poll cycle instead
  of every 600s. Now gates on the more recent of that timestamp and a
  new _last_observe_attempt_ts, stamped on every attempt regardless of
  outcome.
2026-08-06 01:39:03 +00:00
Marc Billow 50bb893407 review: re-arm the settle window on retry, tighten comments, close a test gap
An Opus review of the three prior commits on this branch (PR #306)
turned up two real defects and a documentation/test gap, all fixed
here:

- async_send_command's retry (issue #294) armed the write-settle
  window before the retry existed, so the reconnect pause plus a
  second PUT could eat into the time meant for the confirming poll,
  reviving the revert-then-reapply symptom the window was sized to
  prevent (issue #9). Re-arm it after a successful retry lands.

- test_send_command_reconnects_and_retries_after_socket_closed relied
  on the observe-session fixture's no-op _close_session, so
  self._session never actually went None and _do_put's reconnect
  guard was never exercised -- the test passed even with that guard
  deleted. Now overrides _close_session/_connect_session to actually
  drop and rebuild the session, and asserts the reconnect happened.

- async_send_command's docstring still said "Fire-and-forget", which
  stopped being true the moment it started retrying and raising.

Also trimmed the three comment blocks the review flagged as
reproducing their commit messages verbatim, per CONTRIBUTING.md's
comment-style rules.

One review finding is not addressed here and needs a decision: the
new _close_session() call in the command-retry path isn't
synchronized against _attempt_observe_mode, which touches the session
without _session_lock. A write's reconnect can race an in-flight
observe-mode subscribe attempt and tear down the session it's using.
Fixing it properly means broadening lock scope around observe-mode
entry, which risks blocking a write behind an up to ~15s subscribe
grace period -- a tradeoff not made unilaterally here.

A second finding (dropping the old .strip()'s per-line whitespace
handling in _normalize_pem) did not reproduce against a real
certificate/key, only against the test suite's placeholder PEM body,
so it's left as-is.
2026-08-06 01:06:22 +00:00
Marc Billow 77c2d7831e fix(coordinator): retry a command once after a dead-session reconnect
async_send_command's _do_put caught any exception, logged it, and
returned -- no reconnect, no retry, no error the user could see. A
command landing on a session Samsung's firmware closed between polls
(the same 'known device behavior' _async_update_data already
reconnects around) was silently lost, with nothing to do about it but
a manual reload of the device (issue #294).

Mirror the poll path's own recovery: on failure, close the dead
session, pause, and retry the PUT once against a freshly reconnected
one. If that also fails, raise a HomeAssistantError instead of just
logging, so the user gets a visible error rather than a command that
quietly did nothing. The retry runs under the same session lock the
poll path uses, so a write landing mid-reconnect can't race a
concurrent poll cycle rebuilding the same session.
2026-08-06 00:44:35 +00:00
Marc Billow 252306838d fix(coordinator): downgrade observe mode when a device stays unreachable
When a poll fails and the immediate reconnect retry fails too,
_async_update_data returned the last-known snapshot as a degraded
success (issue #254) without ever touching observe mode. That's fine
for the data itself, but the connection-mode sensor reads straight
from self._observe.mode, and only the *successful* reconnect branch
ever changed it -- so a device that drops off the network entirely
(air-gapped, powered off, Wi-Fi down) left that sensor reporting
"Push" forever, hours after the session was actually dead (issue
#287).

Downgrade to poll mode on the failure branch too, without attempting
an immediate resubscribe: the reconnect that would normally justify
one just proved there's no live session to subscribe on. Recovery
still happens on its own once the device is reachable again, via the
existing poll-mode retry timer (_maybe_retry_observe_mode).
2026-08-06 00:42:31 +00:00
Marc Billow 455ed5b27c fix(config_flow): normalize a pasted PEM before parsing it
A PEM pasted from a text editor can carry bytes cryptography's parser
refuses outright: a UTF-8 BOM some Windows editors silently prepend,
CRLF line endings, and a stray blank line a paste can introduce
between the header/body/footer. None of those are meaningful in PEM,
but any of them surfaces as an opaque InvalidHeader with no hint of
what's wrong -- which is why the same certificate pasted from
Command Prompt's `type` (no BOM, no stray blank lines) loads fine
while the same file opened in an editor and copied doesn't (issue
#291).

Normalize at the point the pasted blob is first captured, not just
before minting the leaf cert: the same string is stored in the config
entry and reused to re-mint the leaf on a future reconfigure, so a
raw copy would keep failing every time it's read back, not just on
the first attempt.
2026-08-06 00:41:43 +00:00
Marc Billow 26e4c9c167 Merge pull request #267 from kkqq9320/fix/air-quality-state-class
fix(air_purifier): record long-term statistics for the particulate sensors
2026-08-05 19:55:56 -04:00
Marc Billow 4e47a1c3d9 Merge pull request #296 from perseus177/ac-good-sleep-halfhours
fix(airconditioner): good_sleep is hours, but the token counts half hours
2026-08-05 08:18:03 -04:00
Marc Billow bea5206c06 Merge pull request #293 from mbillow/claude/ac-filter-reset-cleanup
feat(airconditioner): reset the legacy filter counter locally
2026-08-05 08:16:02 -04:00
perseus177 75e985d392 style: let ruff format the write helper
`ruff format --check` is part of the Validate workflow and my hand-wrapped
version of the dict literal was not what it produces. No behaviour change.
2026-08-05 12:27:24 +02:00
perseus177 93f45cb356 fix(airconditioner): good_sleep is hours, but the token counts half hours
The Sleep_ token was published as if its value were hours. It is not: the
appliance's own app pairs a duration picker with the values it puts on the wire,
one to one, and the pairing is half hours.

  0:00 0:30 1:00 1:30 2:00 2:30 3:00 4:00 5:00 ... 12:00
     0    1    2    3    4    5    6    8   10  ...    24

So the entity capped at 12 hours actually set six, every value asked for was
halved on the appliance, and twelve hours -- the app's own maximum, stated in its
help text -- could not be reached at all. The reading is halved and the write
doubled, and the step drops to 0.5 because that is the resolution the picker
offers.

Half-hour steps are what the app offers below three hours; above that it offers
whole hours only, so a half hour up there is untested rather than known-bad. A
Number cannot change step part-way, and turning this into a Select of the app's
sixteen values would change the entity's domain on every unit that already has
one, so the step stays 0.5 throughout and the comment says why.

The descriptor's own comment used to admit the upper bound was a guess ("only 0
has been observed on hardware"). The guess of 12 was right; the unit it was
expressed in was not.
2026-08-05 12:21:10 +02:00
perseus177 e6909fb847 fix(translations): mirror the new button string into every catalog
test_every_language_mirrors_the_english_catalog is right to fail on this: a key
present only in English falls back silently at runtime, which reads as a
half-finished translation rather than a missing one.

Also fills in the same key for ko.json, which the original fix predates:
Korean's translation catalog was added after this branch was cut and never
picked up filter_time_reset either.
2026-08-05 01:48:08 +00:00
perseus177 9ee0329467 feat(airconditioner): reset the legacy filter counter locally
FilterCleanAlarm_Clear, through the same single-token options merge as every
other setting on /mode/vs/0. Measured on an ARTIK051_KRAC_18K: 2.04 Changed and
FilterTime_95 (9 h 30 min) -> FilterTime_0, still zero on a fresh DTLS session
and on every poll after; none of the other 17 tokens moved and the alarm
entries stayed Deleted.

The counter has had no reset until now, and the descriptor said so: two earlier
rounds against live hardware failed, and the conclusion drawn from them was
that the reset had to be cloud-only. That conclusion was wrong, and the way it
was reached is the interesting part -- it came from diffing every resource the
appliance reports before and after pressing reset in Samsung's app, which
showed only the counter zeroing and the alarm clearing. A trigger token cannot
show up in such a diff, because a trigger is never stored. The appliance's own
app sends this token and skips the write when the counter is already zero.

Both failures stay in the comment, because they say what this is not: writing
FilterTime_0 (the value is not writable -- 5595 -> 5595 after 69 s, 1925 ->
1925 after 65 s, two units, opposite power states), and POSTing the cloud
capability's command name to /actions/vs/0 (real name, wrong transport).

Gated on the FilterTime_ token, so it appears only where there is a counter to
reset; newer boards report filter usage through their own resource and would
need a different mechanism.
2026-08-05 01:46:29 +00:00
Marc Billow e9e278726a Merge pull request #280 from g1za/main
ITA typo fix
2026-08-04 21:40:34 -04:00
Marc Billow ccdfe7088e Merge pull request #281 from atc722/agent/nv9000d-regression-fix
Fix read-only sensor categories and add NV9000D coverage
2026-08-04 21:40:03 -04:00
Marc Billow c423efdd23 Merge pull request #292 from mbillow/claude/code-comments-guidelines-jiv4ta
Add code comment guidelines; dramatically trim excessive comments
2026-08-04 21:36:23 -04:00
Marc Billow 3918b1e5c8 Fill in the Korean translation gaps left by the AI Purify/auto-clean-stop merges
ko.json (PR #283) was written against an older en.json and never picked up
the ten keys two later PRs added: the AI Purify sensing entities
(switch.periodic_air_sensing, switch.periodic_sensing_skip_status,
number.sensing_interval, select.sensing_mode + its three states,
time.sensing_skip_start/end) and button.auto_clean_stop. Merging main
into this branch surfaced the gap via
test_every_language_mirrors_the_english_catalog.

Translated the missing entries, matching the terminology and phrasing
ko.json already uses for adjacent keys (e.g. periodic_air_sensing's
existing binary_sensor entry, auto_clean's "자동 청소"), and kept "AI
Purify" as the untranslated brand name the same way cs/es/it/nl do.
ko.json's topology now matches en.json's exactly; full suite (1213
tests), ruff, and ty all pass.
2026-08-05 01:33:08 +00:00
Marc Billow 6dd4de8b6b Merge remote-tracking branch 'origin/main' into claude/code-comments-guidelines-jiv4ta 2026-08-05 01:32:58 +00:00
Marc Billow 7086b134c0 Add code comment guidelines; dramatically trim excessive comments
CONTRIBUTING.md gains a "Code comments" section: comment the why not
the what, keep it to a sentence or two with a pointer to the load-bearing
evidence, don't re-derive a sibling's already-documented reasoning, and
move failed-attempt investigation logs out of inline comments.

Applied that policy across the codebase: condensed sprawling module
docstrings, per-entity essays, and multi-paragraph rationale blocks down
to their load-bearing conclusions, while preserving the actual "why"
(issue numbers, calibration evidence, gotchas, don't-guess rationale).
No functional code changed — verified via diff review, ruff, ty, and the
full pytest suite (1211 passed).

One inline investigation log (the AC filter-reset "tried and failed"
notes) moved to docs/investigations/ac-filter-reset.md rather than being
deleted, per the new guideline on where that kind of record belongs.
2026-08-05 01:24:17 +00:00
Marc Billow 7f66d21d73 Merge pull request #283 from atc722/agent/korean-translation
Add Korean translation
2026-08-04 21:07:42 -04:00
Marc Billow 9e12992cb4 Merge pull request #284 from rtvanhook/main
Update README.md
2026-08-04 21:01:09 -04:00
Marc Billow b59b5ae1b4 Merge pull request #290 from perseus177/ac-autoclean-stop
feat(airconditioner): stop a running auto clean, and read its progress
2026-08-04 20:53:44 -04:00
Marc Billow f8a7a1fa66 Merge pull request #268 from kkqq9320/feat/avt-ai-purify
feat(air_purifier): expose the AI Purify sensing engine on /airlevelcheck/vs/0
2026-08-04 20:02:16 -04:00
perseus177 290a348017 feat(airconditioner): stop a running auto clean, and read its progress
Three tokens describe the drying cycle these boards run after cooling, and the
switch only covered the first. AutocleanProgress_ is how far a running cycle
has got, and StopAutoClean_ is a channel for ending one early -- its presence
is what says the appliance accepts that at all, which is how the appliance's
own app gates its stop button. Both tokens are reported by the ARTIK051_KRAC_18K
that issue #136 was about, and by every KRAC fixture here.

The percentage scale is the app's own: it renders the token into a
`<progress max="100">` with a "{{value}}%" label beside it. An idle unit reports
1 rather than 0 -- the same floor the laundry firmware's progressPercentage sits
at when Ready -- so 0-vs-1 is not a reliable "is it running" test, and the
button is deliberately not gated on it.

The sensor shares AUTO_CLEAN's catalog entry the way auto_clean_legacy already
shares the switch's: same figure, different board generation, distinct key so
nothing collides if a board ever reported both.

Stacked on #289 (this branch is cut from it) -- rebase or merge that first.
2026-08-05 00:45:20 +02:00
GeekERDr 54a676f841 Update README.md 2026-08-04 05:31:24 -05:00
hoon a38437fca4 Add Korean translation 2026-08-04 17:13:32 +09:00
kkqq9320 15279066b5 review: move the state_class into the shared tuple's fourth column
The frozenset was a parallel structure for a per-row fact, and the comment
above the tuple already described it as a fourth column -- so the comment
promised the right shape and the code did something else. Fixed to the shape
the comment described: _AIR_QUALITY_SENSORS carries state_class per row and
the comprehension unpacks it, with _RECORDED_AIR_QUALITY and its duplicated
rationale block deleted.

air_monitor imports the same rows and now unpacks four, but discards the
fourth. That board (issue #210) has stamped all five readings as
`measurement` since it was added; consuming the column would silently drop
long-term statistics for Odor and CleanLevel on shipped devices, which is a
behaviour change this branch has no evidence to make. The grade/concentration
split stays scoped to the air purifier.

test_shared_sensor_tuple_keeps_its_three_column_shape asserted the premise
this replaces -- that widening the tuple breaks air_monitor's import -- so it
is replaced rather than renumbered: one test that the rows carry their own
state_class, and one that air_monitor still imports and still stamps all five.
2026-08-04 15:25:12 +09:00
hoon 3d0dca20f9 Add NV9000D cooktop coverage and fix sensor setup 2026-08-04 14:47:42 +09:00
kkqq9320 96d06369bc review: floor the sensing interval at one minute
Dropping native_min to 0 fixed the read range and quietly opened a write:
native_min governs what the user can enter, not just what renders, so 0
became enterable and would have gone out as periodicSensingInterval "0".
Nothing establishes what that does to this board -- both fixtures report 600,
the app's smallest choice is 10 min, and 60 s is the lowest value confirmed
accepted. The two precedents leaned on differ in exactly the way that matters:
oven.cook_time and operational.delay_start_hours sit at a zero floor under a
value where 0 is a real setting ("no timer", "no delay").

One minute is also the resolution this board reports results at.
lastSensingTime lands on an exact minute on every sample from the AVT-WW-TP1
and A-VTWW-TP2 boards -- both fixtures, plus eleven consecutive live readings
-- where the TP1X/AC/hood boards report arbitrary seconds. A sub-minute
interval is unobservable here whether or not the board honours it.

So native_min goes to 1 rather than 0, and the read rounds up instead of to
nearest so a sub-minute reading renders as 1 rather than falling below the
entity's own floor. The write still refuses anything under a minute -- a None
return, the silent no-op range_hood._lamp_level_write uses for a level the
device didn't advertise -- since native_min only guards the UI path, not a
service call.
2026-08-04 14:44:54 +09:00
g1za 330b2f344f ITA typo fix 2026-08-04 07:34:53 +02:00
kkqq9320 a5484f746a review: unfold AI Purify into one entity per field
Review feedback on #268. The largest change is that the sensing-mode select
no longer folds two device fields into one control.

periodicSensingActivationState and autoExeState are independent knobs, and the
appliance presents them that way -- its own UI has an on/off for AI Purify
separately from the three mode choices. Folding them lost two things: a
configured action was invisible while the feature was off, and no select
option could toggle the feature without also overwriting the action. The
switch was not the duplicate it looked like.

So the switch now owns periodicSensingActivationState alone, and the select
owns autoExeState alone. That resolves the hardcoded-options finding at the
source rather than working around it: the select reads supportedAutoExeState
via options_field -- the same shape SOUND_MODE already uses for
supportedModes -- instead of carrying a typed-in tuple, so a board advertising
a fourth action is accepted on both the options list and the write path.
_sensing_mode, _sensing_mode_write and _SENSING_MODE_BODIES are all gone with
the fold.

Option slugs are now the advertised values lowercased (off / airpurify /
alarm) rather than invented names. The catalog carries the labels, so the two
'off's stay distinguishable in the UI: the switch's means the unit isn't
sampling, the select's means it samples and doesn't act on the reading -- what
the app calls "sensing only".

Also from the review:

  * _interval_minutes checks `is None` so a reported 0 stays 0, and native_min
    drops to 0 since sub-30s values round there. oven.cook_time and
    operational's delay hours are the precedent -- both convert a device time
    value and floor at zero. The Number-rather-than-Select choice is now
    stated in the write helper: the app offers three fixed intervals, but this
    resource advertises no supported-values or range field (supportedAutoExeState
    sits right beside it, so the board does advertise constraints where it has
    them) and it accepted 60 s, six times finer than the app's smallest choice.
  * _skip_time_write no longer splices a malformed half back onto the wire.
    The read side already refuses one it can't parse; the write side now
    zeroes it to match.
  * air_sensing_state and last_air_sensing_level lose enabled_default=False,
    matching range_hood.AIR_LEVEL_CHECK. Hiding two of three read-only keys
    while claiming key parity with that capability -- and leaving the third
    visible -- had no justification behind it.
  * The catalog-parity test drops periodic_air_sensing from its key set: that
    key is a SwitchDesc here and a BinarySensorDesc on the hood, so the two
    live in different platform catalogs and are worded differently. The claim
    now covers only the three read-only sensor keys, where it holds.
  * Tests route through the descriptors (_desc(key).write_fn / .value_fn)
    rather than module-private helpers, matching test_air_monitor_capabilities.

startSensingOnce stays unbound, now explicitly rather than by omission -- the
module comment records it as deferred. It looks like a one-shot "sense now"
button, but this board acknowledges writes it discards, and nothing has
confirmed the side effect yet.

Goldens are untouched: the key set is unchanged, only sensing_mode's value
moves from the folded slug to the raw autoExeState.
2026-08-04 14:20:58 +09:00
kkqq9320 412fff9b99 feat(air_purifier): expose the AI Purify sensing engine on /airlevelcheck/vs/0
/airlevelcheck/vs/0 has been covered as "periodic air-quality sensing
scheduler plumbing" since the registry gained a coverage stub for it. Two
AVT-WW-TP1-23-AXX500 dumps (issues #84 and #190) show it is not plumbing: it
drives the feature the SmartThings app calls AI Purify, where the unit wakes
on a timer, samples the air, and optionally acts on the result. Every field
is named, none are opaque, and two of them are already user-set on the
reported units.

The select's three on-states are the app's own options rather than an
invented grouping -- it offers exactly "Sensing only" (sample, take no
action), "Auto clean" (purify while the air reads bad, stop once it
improves) and "Get notified" (raise a SmartThings notification). Labels were
transcribed from the Korean app and rendered in English; the auto-stop half
of "Auto clean" is the app's own description and is not otherwise visible in
the dump, which reports only the selected autoExeState. The remaining entity
names follow their raw fields rather than inventing a concept -- the skip
window is "sensing skip", after periodicSensingSkipStatus/Time.

Three of this registry's four board families report the resource with the
same field names -- TP1X_DA-AC-AIR (#130), A-VTWW-TP2 (#151) and AVT-WW-TP1
(#84, #190). Only ARTIK051_TVTL (#56) has no such href, and its golden is
unchanged. Bound unconditionally rather than behind a match_fn; the one field
that genuinely varies (periodicSensingInterval, absent on the #130 board) is
gated per-entity, so that board gets eight entities instead of nine rather
than a broken one.

range_hood.AIR_LEVEL_CHECK already models this same href, and its read-only
keys are reused verbatim here so both families share one catalog entry. It is
deliberately not imported: the hood exposes periodic_air_sensing as a
read-only BinarySensorDesc and this board needs a writable SwitchDesc on that
key, so reusing the hood's capability would migrate every hood user's entity
to a different platform.

Every write was exercised on AVT-WW-TP1-23-AXX500 hardware. This board hands
out 2.04 for writes it silently discards (see HEPA_FILTER's filter-reset
note), so an echo proves nothing -- each was judged by whether the value
survived a reconnect, which forces a new DTLS session, fresh discovery and a
fresh observe of the href, leaving no cached state to read back:

  * sensing_mode's combined two-field PUT lands both fields, both ways:
    sensing_only -> auto_purify raises autoExeState with activation still On,
    and back again lowers it.
  * The sensing-skip switch holds Off -> On and back.
  * The half-preserving time writes hold: from 13:00-23:00, writing
    start=07:30 then end=22:00 left the device on '07302200' -- each write
    kept the half it wasn't given.
  * periodic_air_sensing and sensing_interval: writing 60 s drove an observed
    ~60 s sensing cycle.
  * The read side of the skip window is separately cross-confirmed on two
    units: #84's sits at the inert '00000000', #190's carries a real
    '03002300' (03:00-23:00), which is what pins the HHMMHHMM split.
  * The other two families get the writes on field-shape grounds -- the same
    basis on which they already share MODE, HEPA_FILTER and the air-quality
    sensors.

range_hood._timestamp moves to common.epoch_to_utc so both callers share it,
matching how filter_usage_percent was shared. No behaviour change.

Every existing entity is untouched: the three golden updates are purely
additive, no renames, no unit or device_class changes.
2026-08-04 12:57:03 +09:00
Marc Billow 0c2e219464 Merge pull request #273 from mbillow/claude/device-discovery-config-flow-dmeagp
Rebuild device discovery on the ClientHello probe and resolve identity up front
v0.19.0
2026-08-03 20:36:17 -04:00
Marc Billow b5699badbe Stop the v1 migration re-keying devices onto a placeholder serial
_serial_from_unique_id took the entry's unique_id at face value. That is
right for an entry whose unique_id holds a real serial, but the unique_id
records what the config flow believed when it ran, not what the registry
holds now -- and for two firmware families those are different things.

Entries added before the placeholder rules landed (issues #83/#189) were
keyed on the placeholder itself: `localthings_Nothing(SVC)` for the
ARTIK051_DONGLE_REF dongles, `localthings_FFFFFFFFFFFFFFF` for the
DA_WM_A51_20_COMMON laundry boards. The coordinator has been resolving
those same boards to the host ever since, so their devices and entities
are host-keyed today. Migration read the placeholder back off the
unique_id, decided the host-keyed rows were the stale ones, and rewrote
them onto the placeholder -- reintroducing exactly the collision those
issues exist to prevent, since every unit of the family reports the same
placeholder and would go back to sharing entity unique_ids.

Run the recovered string through resolve_serial, which is the whole point
of that helper being shared. The old `host:port` special case stays: it's
a config-flow-history artifact rather than a device-reported serial, so
resolve_serial can't recognize it.

The repair pass had a second, narrower way to lose data. Removing a device
takes its entities with it (entity_registry.async_device_modified), and
the removal branch ran after the entity pass -- so an entity that had just
been re-keyed rather than removed, because its serial-keyed key was free,
was destroyed a few lines later along with the entity_id, name and area
the rewrite existed to preserve. Move surviving entities onto the device
they now belong to before removing the duplicate.

Reachable when the serial-keyed device exists but a given entity's
serial-keyed key doesn't -- e.g. the user deleted the visible duplicate by
hand, which is the first thing anyone hitting #236 tries.

Also fold the modelNum `<model>|<board>` split into resolve_model beside
resolve_serial. The config flow and _run_discovery each had their own copy
under a comment promising they matched; a device that renames itself on
the first poll is what a drift there looks like.
2026-08-04 00:12:40 +00:00
Marc Billow d5adf311da Normalize cs.json to LF line endings
The Czech catalog was the only file in the repo still using CRLF, which
made every edit to it show up as a whole-file rewrite in diffs and hid the
one line that actually changed.

Content is byte-identical apart from the line endings, and the file now
matches the exact json.dumps(indent=2, ensure_ascii=False) formatting the
other four catalogs already use.
2026-08-03 20:33:30 +00:00
Marc Billow 6033709f24 Replace the blanket "cannot connect" with a real failure taxonomy
Adding a device had one message for nearly every way it could fail: "Cannot
connect to the device. Verify the IP address is reachable and the CA
credentials are correct." That covers an IP with nothing on it, an
appliance on cloud-only firmware, a device still holding the session from
the last attempt, a device that answered and rejected our certificate, and
Home Assistant having no internet to reach Samsung's cloud. Only one of
those is fixed by checking the IP and the CA credentials, and the message
gave no way to tell which one you had.

The probe already gathers enough to tell them apart, so classify it:

- cert_rejected     the appliance sent a certificate alert. The CA
                    credentials aren't the AC14K_M CA it trusts, or they
                    don't pair. Far and away the most common real setup
                    mistake, and previously indistinguishable from a typo
                    in the IP address.
- handshake_failed  a fatal alert unrelated to the certificate (protocol or
                    cipher mismatch) -- no amount of fiddling with CA
                    credentials will fix it.
- handshake_timeout the ClientHello probe proved a DTLS server is right
                    there, but the handshake never finished. Usually the
                    appliance is still holding the association from a
                    previous attempt; it clears on its own in about a
                    minute.
- ports_closed      ICMP port-unreachable on the whole range: something is
                    at that address and it isn't exposing a local API.
                    Cloud-only firmware (TCP 8888 only) lands here.
- no_dtls_server    some ports open|filtered, none speaking DTLS -- likely
                    another device on that IP.
- no_response       nothing came back at all.
- cloud_unreachable couldn't reach Samsung's cloud gateway for the UUID.
                    An internet problem on HA's side, not the appliance's.
- unexpected_response  authenticated fine, then returned something we can't
                    read. Neither connectivity nor credentials.

Certificate alerts are read back out of the error text OpenSSL puts in
DtlsCoapSession's ConnectionError, not by re-probing. The library's
diagnostic probe would report the alert authoritatively, but it drives the
handshake far enough to commit association state on the device -- and an
orphaned association is exactly what makes the *next* attempt time out
(RFC 6347 4.2.8), which is a bad trade on a path the user is about to
retry.

Telling ports_closed from no_response needs the UDP sweep to separate a
refusal from an unreachable. Both leave a port "not live", but ECONNREFUSED
is a *response* -- the host is there -- while EHOSTUNREACH/ENETUNREACH mean
the datagram never left. A wrong IP on the local subnet never answers ARP
and fails every send that way, so treating the two alike would have told
those users their appliance was on cloud-only firmware. The sweep now
returns live/refused/unreachable separately, and the preferred-port rescue
moved out of it into _sweep_ports: the rescue is a candidate-selection
decision, and folding it into the sweep's verdict destroyed the evidence
the message is built from.

Every failure carries the error key that fits it, so the flow maps
exceptions instead of guessing, and logs the specifics (alert name, per-port
outcome, response code) at warning level -- the messages that mention the
log now have something to point at.

Certificate re-minting for a reused leaf is also narrower and more correct
as a result: it now triggers on CertRejected specifically, rather than on
"every attempt raised ConnectionError and a port was confirmed".

All five translation catalogs carry the eight new messages. The non-English
ones are my own work rather than a native speaker's; corrections welcome.
2026-08-03 20:29:53 +00:00
Marc Billow 15be379243 Rebuild device discovery on the ClientHello probe and resolve identity up front
Two problems, one setup path.

Port detection (issue #211): the config flow found the DTLS port by
elimination -- a 1-byte UDP probe can't tell a silent port from a real
DTLS server, so every port it couldn't rule out got a full certificate
handshake, and every false positive cost the whole 12s HANDSHAKE_TIMEOUT_S
before the next was tried. Adding an appliance took 30-40s.

smartthings-local 0.1.2 ships a stateless ClientHello probe that settles
this positively: a real DTLS server answers with a HelloVerifyRequest in
~1 RTT, and per RFC 6347 4.2.1 it does so without allocating association
state, so the probe leaves nothing behind on the appliance. The whole
49152-49160 range is probed at once and exactly one confirmed port is
given a certificate handshake. Fanning out is safe here in a way racing
real handshakes is not -- each probe is bounded by a 3s budget, so the
pool costs one probe's wall clock rather than the sum of the range, with
no losing threads left running behind us.

The UDP sweep stays as the fallback for when the probe confirms nothing:
it errs in the opposite direction (it reports everything it can't rule
out), so it still surfaces a device on a path that eats our ClientHello,
and it keeps its issue #192 preferred-port rescue.

Port detection now runs first and needs no credentials, so an unreachable
host fails before any round trip to Samsung's cloud. And a second
appliance reuses the existing entry's leaf cert rather than re-minting --
every device accepts the same one -- which makes adding one independent
of Samsung-cloud reachability. A confirmed-live device rejecting the
reused leaf (the UUID does rotate) re-mints and retries once, so reuse
stays self-correcting; a timeout doesn't, since a fresh cert can't fix
nothing answering.

Identity (issue #236): the coordinator seeded device_serial with the
configured host and only replaced it after the first successful poll. But
device_serial mints *permanent* registry keys -- entity unique_ids and
device identifiers -- so anything registering before that poll returned
was written into the registry keyed on the IP address forever. The
connection-mode sensor is added unconditionally rather than from `bound`,
so it was the reliable victim: when the serial-keyed identity appeared
moments later HA created a second device and entity, and the IP-keyed
pair was orphaned. Deleting them didn't help; the next restart that lost
the race recreated them.

The probe already learns the identity, so store it on the config entry --
serial, model, manufacturer, device type. The coordinator seeds
device_serial and its DeviceInfo from those at construction, so keys are
correct from the first entity that registers even if the first poll is
slow or fails outright. There is no placeholder left to correct.

Discovery now treats the registered identity as authoritative rather than
re-keying a device that already has registry entries; it adopts and
persists the polled identity only for an entry that has none, and warns
if a different appliance answers on the same IP.

Entry version 1 -> 2 recovers the serial from the entry's unique_id (the
flow has always keyed it on the probe's serial) and repairs what the old
registration orphaned: IP-keyed devices and entities are rewritten in
place where the serial-keyed key is free -- keeping entity_id, name, area
and every automation referencing them -- and removed where both exist,
since the IP-keyed one has been dead since the restart that made it.
Placeholder-serial boards (issues #83/#189) were keyed two ways at once,
`host:port` on the entry and `host` in the registry; migration collapses
the entry onto the registry's form. One resolve_serial() now serves both
sides, so they can't drift apart again.

The remaining step in the desired pipeline -- probe for subdevices, then
register devices, then populate entities -- already holds:
_enumerate_subdevices_blocking runs before _run_discovery, which runs
before platforms are forwarded. Duplicating it in the config flow would
mean re-running Pattern B's per-href fallback probe, which is the
opposite of what issue #211 is about.
2026-08-03 20:14:54 +00:00
Marc Billow cdaff4a1ca Merge pull request #272 from mbillow/claude/issue-triaging-fuq5sn
Issue triage batch: dehumidifier, AC, dryer, fridge, dishwasher, climate fixes
2026-08-03 15:45:12 -04:00
Marc Billow 6a6eef25b8 Fix ty type error in test_dryer_drum_clean.py
for_device_by_model returns DeviceRegistry | None; accessing .capabilities
directly off the inline call result left the None case unnarrowed. Switched
to the same reg/resources-tuple helper pattern every other by-model test
file in this suite already uses, which ty resolves cleanly.
2026-08-03 19:38:57 +00:00
Marc Billow da25d567cb Remove pointless catalog-literal tests; document the anti-pattern
Two tests added while triaging #244/#226 just re-asserted a translation
string against the catalog entry that had been written moments earlier
(dryer_cycle_table_03's '51'/'53'/'4e', dishwasher_cycle's '83'/'86').
Neither exercises any code path -- they pass by construction and only
break when someone later edits the label text for wording, not when the
actual code/value mapping regresses. tests/test_translations.py already
holds the invariants that matter for catalog data.

Documents the anti-pattern in the adding-device-support skill so future
translation-only fixes don't reach for this pattern again.
2026-08-03 19:32:44 +00:00
Marc Billow e89aa4bab5 Fix transposed Normal/Express 60 dishwasher cycle labels (issue #226)
'83' and '86' were swapped in the dishwasher_cycle catalog. Both the
original DW9000F-class fixture this table was built from and the issue
#226 reporter's board share the identical DeviceType_0812, and the
original fixture's own editCourseList puts the two codes back to back
(positions 4-5) -- a plausible adjacent-pair transcription slip. The
reporter's live confirmation (selecting 'Normal' ran the physical Express
60 program and vice versa) settles which way: '86' is Express 60, '83' is
Normal.

The energy-sensor part of the same issue was already resolved per the
issue thread (the device genuinely doesn't report usage, so the sensor's
removal was correct) -- not touched here.
2026-08-03 19:30:02 +00:00
Marc Billow fb5ed32b0f Add discrete freezer setpoint support for TP1X_REF_21K (issue #229)
The reporter's fridge/freezer combo reports the issue #186 discrete
definite-setpoint pattern on both compartments, but only the cooler half
was modeled -- /temperature/definite/freezer/vs/0 was unbound. Adds
DEFINITE_TEMPERATURE_FREEZER, identical shape to the existing cooler
capability (same fields, just negative supportedList values).
2026-08-03 19:24:59 +00:00
Marc Billow 1becd85f6e Fix HOMECARE_WIZARD_V2 false-positive warning and None entity_id log (#235)
HOMECARE_WIZARD_V2 appears in /mode/vs/0's supportedModes on TP2X_RAC_20K
units but is a capability/option flag echoed from
/configuration/vs/0's airconOptionList, not a selectable HVAC mode -- the
unit's current mode never reports it. Added to a new
_NON_HVAC_OPTION_CODES set that's dropped silently, so hvac_mode/hvac_modes
stop tripping the issue #93 unmapped-mode warning for it on every start.

Also fixes _warn_unmapped logging "None: device mode ..." during setup's
first discovery pass, before the entity is added to hass and entity_id is
assigned -- falls back to unique_id (set eagerly in __init__), so multiple
same-type devices are distinguishable in the log.
2026-08-03 19:22:06 +00:00
Marc Billow 0cc9486ad5 Add missing dryer cycle labels for DV90DG6845LHU5 (issue #244)
Codes 51 (Eco Cotton), 53 (AI Dry+), and 4e (Self Dry) were confirmed by
the reporter selecting each program on the physical appliance and reading
back the resulting raw course code, same table (Table_03) as the existing
issue #80 confirmations.

Also fixes an import-sort lint error left over in by_type/dehumidifier.py
and a stale comment in dryer.py claiming codes 21/4c were still
unidentified when the catalog already had them.
2026-08-03 19:17:41 +00:00
Marc Billow e6d7dcddc8 Add dryer Drum Clean+ tracking; fix multi-entry DrumCleanLog parsing (#258)
Dryers report the same DrumCleanProposal_/WashingTimes_/DrumCleanLog_
options[] tokens washer.py already models for issue #9, so
drum_clean_cycles_remaining/drum_clean_last_cleaned move to laundry.py and
get bound on dryer's /course/vs/0 too.

DrumCleanLog_ on the reporter's dump is a '|'-joined history of every past
clean rather than washer's single bare timestamp -- the shared helper now
takes the last (most recent) entry, which turns out to also fix a latent
bug on four existing washer fixtures whose own DrumCleanLog_ was already
multi-entry and silently failing to parse into drum_clean_last_cleaned.

No heat-exchanger-clean tracking was found in either dump #258 supplied;
noted in dryer.py so a future report knows this was checked.
2026-08-03 19:14:29 +00:00
Marc Billow cf09247e39 Add AC UV LED, ventilation alarm, and PM1 filter support (issue #270)
TP1X_FAC_TIME_23K reports three previously unbound hrefs: a UV-C
sterilization LED and a ventilation-reminder alarm (both plain On/Off
toggles), and a second PM1-rated dust filter with no live usage/status
fields on this particular dump.

The PM1 filter capability gates each entity on its own field's presence
rather than a blanket ignore, since the TP1X_DA-AC-CAC-01001_0000 cassette
AC (issue #191) reports the same href with full live data -- this also
closes two of that device's ten documented coverage gaps (UV LED and the
PM1 filter) as a side effect.
2026-08-03 19:07:03 +00:00
Marc Billow 0994ca487a Restore CRLF line endings in cs.json
The previous commit's translation update rewrote this file with LF
endings; every other language file in the catalog already uses LF, but
this one was CRLF before that change.
2026-08-03 19:00:06 +00:00
Marc Billow 3ef64eae52 Add dehumidifier display switch and watertank lighting (issues #271, #231)
The TP1X_DA_AC_DHM_01001_0000 revision (model AY70H18100GTD) additionally
reports /display/vs/0 (same shape as air_purifier's screen toggle, reused
directly) and /watertank/lighting/vs/0 (on/off, color, and brightness for
the tank's ambient light, plus a diagnostic alarm-status flag). Both issues
submitted the identical dump, so one fix covers both reports.

Also adds x.com.st.d.dehumidifier to the /oic/d device-type routing table
now that a dump confirms it.
2026-08-03 18:59:22 +00:00