Commit Graph
360 Commits
Author SHA1 Message Date
Marc Billow 95ce358d55 fridge: fold the three Auto Door Open variant hrefs into one pattern cap (issue #328)
AUTO_DOOR_SINGLE/KIMCHI/WINECELLAR were identical one-line no-entity
Capability declarations differing only by href. Replaced with
AUTO_DOOR_VARIANT, a pattern cap keyed on href_prefix='/autodoor/' and
gated by match_fn (presence of ado.openOptions) rather than the prefix
alone, so it only claims the variant-declaration hrefs and not
/autodoor/timer/vs/0 -- which doesn't matter in practice anyway, since
that href's own exact-href AUTO_DOOR_TIMER cap always wins first.

Registry-scoped (refrigerator.py's own pattern_capabilities list), not
global ignored.py -- the unknown-device-type fallback that motivates
ignored.py's 'exact hrefs only' rule never reaches this registry, so the
same constraint doesn't apply. A fourth fridge sub-type reporting this
feature at a new href now needs no code change to stay covered.
2026-08-09 01:06:37 +00:00
Marc Billow 5e2c23a62d Add device support for dual-cavity range TP1X_DA-KS-RANGE-0101X (issue #324)
This board (NE63T8751SG/AA-class) reports no /information/vs/0 at all --
the modelNum-based routing fallback has nothing to read -- so it fell
back to 'unknown' and lost the whole range registry (oven mode/setpoint/
door/connected, cooktop monitoring). /oic/d does carry oic.d.range,
though, so this is a routing fix, not a new capability: adds 'oic.d.range'
to _OIC_TYPE_TO_KEY.

The second oven cavity is a genuine Pattern A indexed subdevice at
/device/1 (issue #177's mechanism) -- once routing resolves the master to
the range registry, the same registry already applies to the subdevice's
canonical view and every href on both binds with zero gaps.

_discover_full gains an optional device_types param (default (), every
other fixture unaffected) so a fixture that can only route via /oic/d can
exercise the same subdevice-aware pipeline the other composite fixtures
already do.
2026-08-09 00:58:26 +00:00
Marc Billow b6f0bc22cb Add device support for Samsung Refrigerator TP1X_REF_21K auto-door variants (issue #328)
Three new dumps from one household's TP1X_REF_21K fleet (regular
single-door, kimchi, wine cellar) exposed the Auto Door Open feature's
timer and voice/sound feedback toggles, plus wine-cellar-specific
coverage: a deodorizing filter at its own href, a multi-compartment
pantry select, and a table-revision info resource.

- STATUS_LOCK gains auto_door_voice_control/auto_door_sound_control,
  gated on each field's own presence.
- New AUTO_DOOR_TIMER (a discrete-options select, same shape as the
  DEFINITE_TEMPERATURE_COOLER/FREEZER pattern) and three no-entity
  AUTO_DOOR_SINGLE/KIMCHI/WINECELLAR coverage hrefs -- every dump seen
  reports exactly one openOptions value with no paired current/desired
  field to choose against.
- New DEODOR_FILTER (reuses AIR_FILTER's entities at a different href),
  WINECELLAR_PANTRY_ZONE, and WINECELLAR_INFO.
- by_type: oic.d.krefrigerator and x.com.st.d.winecellar routed to the
  refrigerator registry via /oic/d, alongside the existing modelNum-based
  routing.
- kimchi_zone_mode's translation catalog gains three supportMode codes
  (bare storage_fridge/storage_freezer without the _normal suffix, and
  the apparently-placeholder newmode_kimchi_0000) surfaced by the kimchi
  fixture, across all seven languages.

Three new scrubbed fixtures + goldens + tests, one per variant.
2026-08-09 00:57:36 +00:00
Marc Billow 07c20e82aa Stop double-converting filterUsage on AIR_FILTER/HEPA_FILTER (#330)
filterUsage is already a 0-100 percentage on every confirmed family,
including ARTIK051_PRAC: filterStatus flips to 'wash' at
filterUsage == '100' regardless of filterCapacity (60/224/500 across
other fixtures), which only holds if filterUsage is already a percent.
filter_usage_percent() divided by filterCapacity again, reading a
filter due for washing as 20% fresh.

air_filter_usage_hours had the mirror problem: it read filterUsage
directly as an hour count with device_class=duration, when the field
is a percent. It's now derived from the percentage and filterCapacity
(new filter_usage_hours() in common.py) instead.
2026-08-09 00:37:09 +00:00
Marc Billow f5e99d71e3 Address PR #251 review feedback: no invented English fallback text
washer_cycle_fallback no longer wraps an unrecognized code in an
'Unknown (0xNN)' label -- that baked untranslatable English into a
component built to be fully translatable. It now only ever surfaces a
device-provided personal-course name; an unrecognized standard code
displays as its raw value, same as before PR #251.

Also translates nl.json's '69'/'88' washer labels left in English (same
review), and makes de.json's own 'smart' states consistent with the
'Intelligente Lüftung' translation already used for smartventilation.
2026-08-08 23:44:54 +00:00
Marc Billow 8ed3d1467f Apply ruff format to code merged from PR #251/#275
PR #251 and PR #275 predate this repo's ruff-format adoption on those
files; running the formatter (single->double quotes, line wrapping,
trailing-comma cleanup) keeps the merged code consistent with the rest
of the codebase. No behavior change.
2026-08-08 21:24:02 +00:00
Marc Billowandgalaxysj 00db7890f5 Squash-merge PR #251: fix appliance course labels and AC setup timeouts
- Add confirmed Samsung Table_02 washer course mappings (69-79, 88)
- Add confirmed dishwasher course mappings (82, 8a, a7, a8, 8c, 88)
- Localize new washer/dishwasher course labels in en, cs, nl
- Decode device-provided personal washer course names from TLV payloads
- Show unrecognized washer enum bytes as 'Unknown (0xNN)'
- Normalize select current-state and options through one display path
- Bound first-setup subdevice enumeration with a shared time budget

Co-authored-by: galaxysj <224385302+galaxysj@users.noreply.github.com>
2026-08-08 21:17:43 +00:00
Marc Billow 3675d8087b Scope learning to the device, and simplify the store
The href alone was not a sufficient key. /mode/convenient/vs/0 is
declared by three family registries with three meanings: a real preset
resource on the AC, explicitly unmodeled on the dehumidifier (no live
current-value field), empty on the air purifier. Matching on the href
globally meant a dehumidifier reporting a mode there would learn it,
persist it, and show it in diagnostics for a resource nothing offers.
The coordinator now narrows LEARNABLE to the hrefs a climate entity is
actually bound to, at discovery -- which also retires the per-rep
subdevice walk, since those hrefs are already actual.

With that, LEARNABLE is a plain frozenset of hrefs and the per-href
LearnRule goes away: its two fields were the same module constants for
its only entry. observe() now returns the codes it learned rather than a
bool the caller re-reads the store to interpret, so the log names what
was new instead of everything ever learned.

learned.py also takes ownership of the entry key and persisted shape --
the options flow was the second module that knew both, and the shape has
already changed once.

Comment trims throughout, per CONTRIBUTING: the LEARNABLE entry no
longer recounts how many reporters there were, and three copies of the
same test-stub comment are gone.
2026-08-08 19:49:25 +00:00
Marc Billow 4f3bdde6e5 Address review findings on the learned-modes store
Flatten the store to {href: [codes]}. One href carries one LEARNABLE
rule, so keying the codes by the rule's supported field too let the
write side (rule.supported_field) and both read sides (the module-level
SUPPORTED_FIELD) disagree the moment a rule used a different field --
codes learned and persisted, then never offered.

The options flow's reset step read the persisted value raw in the
entry-not-loaded branch, so malformed data aborted the one screen that
can clear it; route it through LearnedModes like every other reader.
For the same reason forget_learned_modes() now persists whenever the
entry carries a record, not only when the in-memory store had one: a
record _coerce rejected at startup exists only on the entry.
2026-08-08 19:42:20 +00:00
Marc Billow d65735ac47 Remember modes a device reports but never advertises (issue #327)
Some firmware reports a current mode that is missing from the same
resource's supportedModes. An ARTIK051 air conditioner sits in Quiet
while advertising only [Off, Sleep, Speed, Nano, NanoSleep], so HA
showed preset_mode: quiet and then refused to select it. A second
reporter has three identical units where only the two sharing an
outdoor unit hide it, which rules out a real capability difference.

learned.py remembers any such code and the coordinator persists it on
the config entry, so a mode the device only names while it is active
survives a restart. climate._supported unions it into the resource's
own list, which fixes the read and the write together --
async_set_preset_mode reverse-resolves the device code from that same
list.

Learning is allowlisted per canonical href rather than global. Across
the fixture corpus 17 dumps already report a current mode that is not
in supportedModes: an oven idling in NoOperation, a fridge's
/mode/vs/0 carrying capability tokens like WATERFILTER_DISABLE. Those
are not selectable options, and remembering one permanently would put
an option in the UI that the device can only reject. Only
/mode/convenient/vs/0 is learnable today.

On by default, with a per-device option that stops offering and
learning at once, and a reset step in the options flow for a code that
turns out to be bogus. Diagnostics report what was learned separately
from `resources`, which stays exactly what the device said.
2026-08-08 19:10:53 +00:00
Marc Billow 6ee60beae9 Make holding the session across a sequence the caller's choice
Holding _session_lock for a whole write sequence buys certainty about what
the appliance saw and when, but blocks every poll and entity write for the
sequence's full length -- up to 10 x 30s. Which of those matters more
depends on what is being probed, so it is now hold_session_lock on
async_raw_write_sequence and a field on the service, defaulting to the
holding behavior that shipped.

Off, the lock is taken per write and released across the settle waits, so
entities keep updating through a long sequence. Exactly one of the two
context managers is ever the real lock -- asyncio.Lock isn't reentrant.

Tests assert the lock's actual state during the settle wait in both modes,
rather than just that the flag is accepted.
2026-08-07 23:46:28 +00:00
Marc Billow a6d818dfc0 Fix four review findings in the raw write/read services
- services.py: normalize an href before handing it to Subdevice.to_actual.
  That transform is textual and rewrites only a trailing '0' segment, so
  '/mode/vs/0/' passed through it untouched and normalized downstream to
  the master's '/mode/vs/0' -- landing the write on the wrong oven cavity
  while still answering 2.04, with nothing in the response to give it
  away. Same order now on the read path.
- services.py: key `verified` off those same normalized canonicals. It was
  built from un-normalized to_actual output against the coordinator's
  normalized hrefs, so a non-canonical input missed the lookup and handed
  back actual hrefs where the documented contract promises canonical ones.
- coordinator.py: report `held: None` when the verify re-read itself
  didn't come back. A non-2.05 yields an empty rep, against which every
  payload comparison is False, so a 4.04 or dropped read was reported as
  `held: false` -- indistinguishable from the board reverting the write,
  which is the one distinction verify_after exists to draw.
- coordinator.py: on a mid-sequence failure, say how many writes landed
  and which, and still kick the refresh. Raising bare threw that away, and
  the appliance is left holding a partial sequence.

Also documents why `settle` waits inside the session lock while
verify_after's wait deliberately doesn't: a poll landing between two
writes is exactly what the sequence exists to rule out, and the caps
bound the worst case at 10 x 30s.
2026-08-07 22:58:37 +00:00
Marc Billow 5dbe990c1d Add write_resource/read_resource services for probing write contracts (issue #300)
The options-flow "Debug write" panel could only ever do one write to one
href per pass -- not enough for the issue #300 wall oven, whose board
discards settings writes while idle and only keeps them once a cycle is
already running. Finding what starts a cycle needs an ordered sequence of
writes across resources, with real settle delays between them, and a way
to check afterward whether anything actually held.

- coordinator.py: async_raw_write_sequence owns a whole ordered sequence
  under one _session_lock hold (so a poll can't interleave mid-sequence),
  with per-step settle and an optional delayed verify_after re-read done
  outside the lock. async_raw_write is now a one-item wrapper over it, so
  tests/test_coordinator_raw_write.py keeps passing unmodified. Also adds
  async_raw_read, a live GET bypassing the cache -- staleness is exactly
  what makes revert-testing unreliable.
- services.py (new): the two HA services. Device-target resolution scans
  loaded coordinators' MAIN/subdevice identifiers and requires exactly one
  match, so an area/label target can't silently fan a raw write out across
  several appliances. Canonical->actual href translation happens here, not
  in the coordinator, which stays subdevice-agnostic.
- services.yaml (new): selectors/descriptions for both services, inline
  per HA's custom-integration support -- keeps translations/en.json's
  mirror test (test_translations.py) green without touching all 6
  languages for a services block. New exception keys (write caps, device
  target resolution) still went into translations/*.json's existing
  exceptions section, mirrored across all 6 languages.
- __init__.py: adds async_setup to register the services once, process-wide.
- config_flow.py: the debug panel's async_step_debug_edit now calls
  write_resource instead of coord.async_raw_write directly, so there is
  exactly one code path that performs a raw write.
- README.md: new Part 5 documenting both services, with a worked
  write_resource example; points the capability-gap section at them.

tests/test_services.py (new): sequencing/ordering, settle timing, changed
vs. held (the reverted case is issue #300's own symptom), exactly-one-
device resolution, subdevice href translation, validation caps, and the
options-flow panel end to end through the service.
2026-08-07 22:08:28 +00:00
Marc Billow a7dc1db8ff Extract usable parts of PR #316 (System Fresh Air Ventilator support)
PR #316 (fork stale by several months, most of its ~2200-line diff was drift
against main rather than real changes) proposed device support for the
Samsung System Fresh Air Ventilator (ACA-KR-TP2-21-AN9000). Extracted what
holds up, adapted to this project's conventions, and left out what doesn't:

Extracted:
- ventilation_mode select on CLIMATE's own href, gated via
  _is_ventilation_mode_device so it can only ever bind on a device whose
  entire supportedModes set is Purification/Ventilation/SmartVentilation --
  verified against every real AC fixture in the corpus to confirm it can't
  false-positive on an actual air conditioner's climate card.
- WINDFREE / WINDSLEEP switches on their own dedicated hrefs.
- A CO2 sensor on AIR_QUALITY, matching air_monitor.SENSORS' already-bound
  device_class='carbon_dioxide'/unit='ppm' descriptor for the same field
  shape rather than guessing fresh.
- HEPA_FILTER / DEVICE_ACTIVE reuse from air_purifier.py.
- Removing /airlevelcheck/vs/0 from _AC_IGNORED and binding
  air_purifier.AIR_LEVEL_CHECK in its place: the PR's claim that this
  project's old "scheduler plumbing" description was wrong turned out to
  be independently verifiable against two of our own existing fixtures
  (airconditioner_cac and airconditioner_tp1x_da_ac_rac_01011 both already
  carry real, populated periodicSensingActivationState/autoExeState
  values), so this benefits existing users, not just the one new device.

Left out:
- Unit/device_class ('ug/m3', pm10/pm25/pm1) on the existing dust/
  fine_dust/super_fine_dust sensors, sourced from an unverified third-party
  screenshot description. air_monitor.py already has an explicit, reasoned
  rejection of this exact mapping for the exact same three fields:
  Samsung's PM10/PM2.5 convention doesn't confirm where a third tier or a
  PM1 reading fits, and a wrong guess mislabels the reading forever.
- A standalone common.POWER switch -- contradicts this registry's own
  documented design (power is deliberately the climate entity's job) and
  would affect every AC user, not just this device.
- Promoting wind/swing to independent selects for every AC user -- a UX
  opinion, not a coverage necessity, and out of scope for this device's
  own support.
- A model-name diagnostic sensor -- /information/vs/0 is already covered
  via the global ignore list, so this wasn't closing an actual gap.

No raw diagnostics dump for this model was ever attached to PR #316, so
there's no fixture for it here (fabricating one would violate this
project's fixture-integrity rule) -- see
tests/test_airconditioner_ventilation_windfree.py's module docstring.
2026-08-07 16:38:09 +00:00
Marc Billow 8551974719 Fix ty type-check failures in new test files
resolve()/for_device_by_model() return DeviceRegistry | None; four new
test files used reg.capabilities/reg.pattern_capabilities without
narrowing away None first. Add the same 'assert reg is not None' idiom
test_dehumidifier_tp1x_dhm01001_capabilities.py already uses.

Verified against a clean venv running the exact CI commands (ruff format
--check, ruff check, ty check, pytest) rather than trusting a stale local
venv that had picked up a mismatched python3.11/3.13 site-packages split.
2026-08-07 15:52:46 +00:00
Marc Billow e2dcc75ed5 Address Opus review findings on the device-support commits above
- Fix a real bug: airconditioner.SOUND_MODE had no exists_fn, so on
  boards (issue #319's FAC) that never report a live 'mode' value,
  entity.py's default field-presence gate silently kept the select from
  ever registering in HA -- while adapter.flatten() (what the golden/tests
  read) has no such gate, so the tests passed while documenting behavior
  the opposite of what shipped. Gate on supportedModes' presence instead.
- Add airconditioner.MDS_ABSENCE_CLEAN for the CAC-class board's
  /mds/absenceclean/vs/0 -- byte-identical shape to issue #319's
  /csi/absenceclean/vs/0, confirmed rather than guessed, closing one more
  of that board's documented coverage-gap hrefs.
- Add missing translation state labels (all 6 languages) for
  edge_lighting_mode/edge_lighting_color/indicator_light_mode's raw device
  codes, so they render as real words instead of a raw '3000K' -> '3000 K'
  fallback.
- Fix an orphaned comment above SOUND_MODE that actually described the
  unrelated DISPLAY reuse, and correct two inaccurate rationale comments:
  the sound/voice ignore reason claimed a distinction from SOUND_MODE that
  this same dump contradicts, and the /csi/* ignore block's 'same
  reasoning as air_purifier.COVERAGE' precedent only actually covers 1 of
  its 5 hrefs.
- Correct the false 'no board-token match' claim in the FAC test file and
  golden-regression docstring -- 'FAC' is a real _BOARD_TOKEN_TO_KEY entry
  (for_device_by_model alone already resolves this board); add a test
  that actually exercises that path, which nothing previously did despite
  the docstring's claim.
- Drop a tautological burner-slot test that only re-asserted what the
  golden regression test already covers via the same code path.
2026-08-07 14:50:24 +00:00
Marc Billow c203bd42c5 Add edge-lighting and indicator-light support for TP1X_DA-AC-CAC-01001 (issue #288)
Six System A/C cassette units on the same board test_airconditioner_cac.py
already documented as having an incomplete coverage gap gave real dump
evidence for two of its remaining unbound hrefs:

- /edgelighting/vs/0: an accent-light strip with on/off, a Smart/High/Low
  mode, and a Kelvin color-temperature select (3000K/4000K/6500K), all read
  from the device's own live supported-value lists.
- /light/stateful/vs/0: a second, distinct light resource with its own
  on/off and Smart/Low/High mode -- not to be confused with EDGE_LIGHTING
  or DISPLAY_LIGHT's ambient mood light.

convenientMode/operatingOption on /edgelighting/vs/0 stay unexposed: present
on every dump but no evidence of what either actually controls.

Only three hrefs remain in test_airconditioner_cac.py's documented gap now
(absence-clean, sound-optimization, smart-sensing-cooling).
2026-08-07 14:27:25 +00:00
Marc Billow 42fd9c2aa4 Close coverage gap and fix phantom lamp switch for TP2X_DA-KS-WALLOVEN (issue #300)
/diagnosis/vs/0 was the dump's only unbound href, now covered via
dishwasher.DIAGNOSIS (same shape already reused by airconditioner.py).

This steam-oven-class board's /mode/vs/0 options[] carries no UpperLamp_
token at all, unlike the NV7000BS-class board LAMP was proven against --
LAMP had no exists_fn, so it registered anyway, always read Off, and any
write to it was a no-op the device had no reason to honor. Gives it the
same options-token exists_fn gate issue #183 already added to
fast_preheat/natural_steam/energy_saving/cooktop_on_alert.
2026-08-07 14:22:32 +00:00
Marc Billow df5b704f3e Close gas-cooktop coverage gap for TP2X_DA-KS-COOKTOP-000001 (issue #314)
/alarms/vs/0 and /kidslock/vs/0 were the dump's two unbound hrefs -- both
are the exact shapes common.UNIVERSAL already models elsewhere
(common.ALARMS, common.KIDS_LOCK_VS_FALLBACK), picked individually rather
than pulling in all of UNIVERSAL to match this registry's existing
hand-picked-common style.

The six-vs-three burner count the reporter originally asked about is
expected behavior (the board's own /mode/vs/0 options genuinely advertise
six OperationState slots on hardware with three physical burners, with no
per-device signal to tell real slots from phantom ones) -- already
explained on the issue; this commit is scoped to the coverage warning.
2026-08-07 14:18:08 +00:00
Marc Billow 26c9168fb7 Add internal air-filter support for TP1X_REF_21K refrigerators (issue #318)
/filter/airdustfilter/vs/0 was the dump's only unbound href -- this board's
internal deodorizing filter, same filterUsage/filterStatus field pair as
common.WATER_FILTER, but filterUsage here is already a 0-100 percentage
with no filterCapacity to divide by (confirmed by filterStatus=="wash" at
filterUsage=="100"). Uses air_-prefixed keys so a fridge with both a
water and an air filter gets two distinct entities.
2026-08-07 14:13:38 +00:00
Marc Billow edf77309ba Add device support for AILP_DA-AC-FAC-02011 air conditioner (issue #319)
This board routes purely via /oic/d's oic.d.airconditioner type (no
board-token match) and reports several resources the sibling
TP1X_DA-AC-CAC-01001 board (issue #191) left as a documented gap:

- /display/vs/0, /settings/sound/output/vs/0, /settings/sound/volume/vs/0
  now reuse air_purifier.py's identical-shape capabilities instead of
  duplicating them.
- /settings/sound/mode/vs/0 gets a new airconditioner.SOUND_MODE reading
  the live supportedModes field, sharing laundry.py's existing
  voice/tone/mute translation catalog since the value vocabulary matches.
- /csi/absenceclean/vs/0 and /csi/energysaving/vs/0 are new, genuinely
  useful controls (absence auto-clean toggle, energy-saving mode select
  plus its state/operatingStatus diagnostics).
- /dnd/autosleep/vs/0, /outdoorsharing/vs/0, /lifestyle/survey/vs/0,
  /settings/sound/voice/vs/0 and /csi/information/vs/0 are ignored as
  plumbing/unconfirmed data with no user-actionable state.

Also fixes a latent bug in air_purifier.SOUND_VOLUME: boards that report
minLevel/resolution but no maxLevel (this one) would have produced a
min=0/max=0 number entity instead of self-gating off.

Updates test_airconditioner_cac.py's documented coverage gap now that
sound_mode/sound_output/sound_volume are covered there too.
2026-08-07 14:10:13 +00:00
Marc Billow dd953b8150 Merge pull request #304 from perseus177/ac-presets-per-hvac-mode
feat(climate): derive legacy AC presets from the unit's own capability bits, per HVAC mode
2026-08-06 08:46:55 -04:00
perseus177 eed04faaed fix(climate): derive presets only when the board publishes both capability maps
One map is not enough to judge by: with only OptionCode present, every
eoc-gated rule reads None, and None means the board does not publish the map
rather than that the feature is absent. artik051_dongle_fac_18k is exactly that
board and lost WindFree in every mode. Requiring both also keeps these bit
positions inside the family they were documented for -- the FAC and CAC dumps
carry only the older map, with values small enough that RAC positions read as
zeros.

Also from review: an unknown HVAC mode falls back the same way, Comfort is
spelled like the identical Speed rule, the unreachable AIComfort branch is
gone, the Cool code comes from the unit's own supportedModes, DlightCool gains
its catalog entry, and the Single User claim is dropped -- the app's own Single
User command sends Comode_Smart, so there is no distinct token to write.
2026-08-06 12:04:46 +02:00
Marc Billow e3e7f4f43c test: suppress ty's invalid-assignment on the fake-session swap
coordinator is explicitly typed as LocalThingsCoordinator here, so ty
correctly sees _session's declared type (DtlsCoapSession | None) and
flags assigning a FakeObserveSession to it. The fixture's own
_connect_session replacement does the same swap without tripping ty,
but only because its self parameter is unannotated -- ty has nothing to
check the assignment against there. Deliberate here (this is the whole
point of the test: substitute a stand-in session), so silenced rather
than restructured; ty's --add-ignore confirmed the comment syntax
(ty: ignore[...], not the mypy-style type: ignore[...] used elsewhere
in this suite, which ty doesn't appear to honor for this rule).
2026-08-06 02:07:39 +00:00
Marc Billow a3cc918343 fix(coordinator): two gaps a follow-up Opus review found in the split
A second review of the observe-mode phase split (previous commit) found
two real regressions it introduced, both in the same failure family it
was built to close:

- async_send_command's failed-retry branch closed the session, then
  raised without downgrading observe mode -- the downgrade only ran on
  the retry's success path. A retry that also fails still leaves the
  session dead, so mode was left claiming "Push" on a session that no
  longer exists, same as the bug this whole fix targets. Moved the
  downgrade to run right after the close, unconditionally on how the
  retry goes.

- _attempt_observe_mode's stale-session abandon (the identity-check
  branch added in the previous commit) didn't flag a resubscribe. A
  session swap discovered there means a fresh, never-tried session now
  exists, but _last_observe_attempt_ts was already stamped for the
  now-abandoned attempt -- so that new session sat unsubscribed for up
  to _RECOVERY_RETRY_S (600s) instead of being retried on the next
  cycle. Now sets _resubscribe_due, same as the two reconnect paths do.

Also closes two test-coverage gaps the same review surfaced by mutation
testing: no test asserted the lock actually holds during the subscribe
burst (only that it's released for the wait), and no test distinguished
the max() in _maybe_retry_observe_mode's throttle from using
_last_observe_attempt_ts alone -- both mutations left the full suite
green. Added one test for each, plus extended two existing tests for the
bug fixes above; all four confirmed via mutation testing (revert the
fix, watch the new/extended test fail; restore it, watch it pass).

One finding from the same review is intentionally left open: async_close
is the one self._session writer that doesn't take _session_lock, so a
close racing _attempt_observe_mode isn't covered by today's identity
check. This is pre-existing (the lock didn't cover any of
_attempt_observe_mode before this branch's earlier commits either), not
a regression from this branch's work, and is a shutdown/unload-path
question rather than the write-vs-observe-mode race this branch set out
to fix.
2026-08-06 02:01:30 +00:00
Marc Billow 69f93be4dc fix(coordinator): close the observe-mode race an Opus design review found
The command-retry fix (issue #294) added a self._close_session() call to
async_send_command that isn't synchronized against _attempt_observe_mode,
which reads self._session and subscribes to it without holding
_session_lock. A write's retry racing an in-flight subscribe attempt
could tear down the session mid-subscribe -- or worse, land the close
*after* the attempt's grace wait already succeeded, letting it commit
observe mode against a session that's already gone: mode claims "Push"
forever, with nothing left to notice the underlying socket is dead.

Split ObserveManager.try_enter_observe_mode into four pieces
(subscribe_hrefs / await_observe_notifies / enter_observe_mode /
abandon_observe_attempt), keeping try_enter_observe_mode as a thin
wrapper so its direct callers in test_observe.py are unaffected.
_attempt_observe_mode now holds _session_lock only for the subscribe
burst (each send is fire-and-forget, not a network round trip) and
re-checks self._session is sess under the lock right before committing
-- sess keeps the old session object alive, so identity can't be
recycled onto a new one, which is what makes the check sufficient
without a separate generation counter. The wait itself stays lock-free,
so a command write is never blocked behind it.

Two more bugs the same investigation turned up, fixed in the same pass
since they're direct consequences of the design above:

- async_send_command's own successful reconnect didn't downgrade observe
  mode the way the poll path's reconnect already does, leaving the same
  stale-commit problem reachable with zero concurrency at all -- just a
  write's retry succeeding while mode was observe. Replaced the poll
  path's local just_downgraded_from_observe with an instance flag both
  reconnect sites set, so either one triggers an immediate resubscribe.

- _maybe_retry_observe_mode's 600s throttle gated solely on
  last_mode_change_ts, which _set_mode only stamps on an actual
  transition -- a device that never succeeds at observe mode leaves that
  timestamp stuck at construction time, so the throttle opens once and
  never closes again, re-attempting on every single poll cycle instead
  of every 600s. Now gates on the more recent of that timestamp and a
  new _last_observe_attempt_ts, stamped on every attempt regardless of
  outcome.
2026-08-06 01:39:03 +00:00
Marc Billow 50bb893407 review: re-arm the settle window on retry, tighten comments, close a test gap
An Opus review of the three prior commits on this branch (PR #306)
turned up two real defects and a documentation/test gap, all fixed
here:

- async_send_command's retry (issue #294) armed the write-settle
  window before the retry existed, so the reconnect pause plus a
  second PUT could eat into the time meant for the confirming poll,
  reviving the revert-then-reapply symptom the window was sized to
  prevent (issue #9). Re-arm it after a successful retry lands.

- test_send_command_reconnects_and_retries_after_socket_closed relied
  on the observe-session fixture's no-op _close_session, so
  self._session never actually went None and _do_put's reconnect
  guard was never exercised -- the test passed even with that guard
  deleted. Now overrides _close_session/_connect_session to actually
  drop and rebuild the session, and asserts the reconnect happened.

- async_send_command's docstring still said "Fire-and-forget", which
  stopped being true the moment it started retrying and raising.

Also trimmed the three comment blocks the review flagged as
reproducing their commit messages verbatim, per CONTRIBUTING.md's
comment-style rules.

One review finding is not addressed here and needs a decision: the
new _close_session() call in the command-retry path isn't
synchronized against _attempt_observe_mode, which touches the session
without _session_lock. A write's reconnect can race an in-flight
observe-mode subscribe attempt and tear down the session it's using.
Fixing it properly means broadening lock scope around observe-mode
entry, which risks blocking a write behind an up to ~15s subscribe
grace period -- a tradeoff not made unilaterally here.

A second finding (dropping the old .strip()'s per-line whitespace
handling in _normalize_pem) did not reproduce against a real
certificate/key, only against the test suite's placeholder PEM body,
so it's left as-is.
2026-08-06 01:06:22 +00:00
Marc Billow 77c2d7831e fix(coordinator): retry a command once after a dead-session reconnect
async_send_command's _do_put caught any exception, logged it, and
returned -- no reconnect, no retry, no error the user could see. A
command landing on a session Samsung's firmware closed between polls
(the same 'known device behavior' _async_update_data already
reconnects around) was silently lost, with nothing to do about it but
a manual reload of the device (issue #294).

Mirror the poll path's own recovery: on failure, close the dead
session, pause, and retry the PUT once against a freshly reconnected
one. If that also fails, raise a HomeAssistantError instead of just
logging, so the user gets a visible error rather than a command that
quietly did nothing. The retry runs under the same session lock the
poll path uses, so a write landing mid-reconnect can't race a
concurrent poll cycle rebuilding the same session.
2026-08-06 00:44:35 +00:00
Marc Billow 252306838d fix(coordinator): downgrade observe mode when a device stays unreachable
When a poll fails and the immediate reconnect retry fails too,
_async_update_data returned the last-known snapshot as a degraded
success (issue #254) without ever touching observe mode. That's fine
for the data itself, but the connection-mode sensor reads straight
from self._observe.mode, and only the *successful* reconnect branch
ever changed it -- so a device that drops off the network entirely
(air-gapped, powered off, Wi-Fi down) left that sensor reporting
"Push" forever, hours after the session was actually dead (issue
#287).

Downgrade to poll mode on the failure branch too, without attempting
an immediate resubscribe: the reconnect that would normally justify
one just proved there's no live session to subscribe on. Recovery
still happens on its own once the device is reachable again, via the
existing poll-mode retry timer (_maybe_retry_observe_mode).
2026-08-06 00:42:31 +00:00
Marc Billow 455ed5b27c fix(config_flow): normalize a pasted PEM before parsing it
A PEM pasted from a text editor can carry bytes cryptography's parser
refuses outright: a UTF-8 BOM some Windows editors silently prepend,
CRLF line endings, and a stray blank line a paste can introduce
between the header/body/footer. None of those are meaningful in PEM,
but any of them surfaces as an opaque InvalidHeader with no hint of
what's wrong -- which is why the same certificate pasted from
Command Prompt's `type` (no BOM, no stray blank lines) loads fine
while the same file opened in an editor and copied doesn't (issue
#291).

Normalize at the point the pasted blob is first captured, not just
before minting the leaf cert: the same string is stored in the config
entry and reused to re-mint the leaf on a future reconfigure, so a
raw copy would keep failing every time it's read back, not just on
the first attempt.
2026-08-06 00:41:43 +00:00
Marc Billow 26e4c9c167 Merge pull request #267 from kkqq9320/fix/air-quality-state-class
fix(air_purifier): record long-term statistics for the particulate sensors
2026-08-05 19:55:56 -04:00
perseus177 2d772f16a7 feat(climate): offer legacy presets per HVAC mode, from the unit's own capability bits
The fixed list of six was offered in every mode on every legacy board. The
appliance publishes what it has as two bit maps in /mode/vs/0's options, and its
own app gates each comfort mode on a bit plus the current mode; this transcribes
that logic. WindFree also needs the mode written before it in Auto, which is
measured rather than assumed.
2026-08-05 17:24:40 +02:00
perseus177 30bd0fd2af fix(airconditioner): Good Sleep needs the mode token its duration belongs to
Sleep_<n> written on its own is answered 2.04 Changed and then discarded, so
the Number wrote nothing at all. Nano wind shares the same Comode_ slot, which
is why writing the nano preset over a running timer silently changed its
duration, and why the two sleep codes the board reports had to become presets:
a preset_mode outside preset_modes is not a state HA allows.
2026-08-05 15:46:53 +02:00
Marc Billow 4e47a1c3d9 Merge pull request #296 from perseus177/ac-good-sleep-halfhours
fix(airconditioner): good_sleep is hours, but the token counts half hours
2026-08-05 08:18:03 -04:00
perseus177 93f45cb356 fix(airconditioner): good_sleep is hours, but the token counts half hours
The Sleep_ token was published as if its value were hours. It is not: the
appliance's own app pairs a duration picker with the values it puts on the wire,
one to one, and the pairing is half hours.

  0:00 0:30 1:00 1:30 2:00 2:30 3:00 4:00 5:00 ... 12:00
     0    1    2    3    4    5    6    8   10  ...    24

So the entity capped at 12 hours actually set six, every value asked for was
halved on the appliance, and twelve hours -- the app's own maximum, stated in its
help text -- could not be reached at all. The reading is halved and the write
doubled, and the step drops to 0.5 because that is the resolution the picker
offers.

Half-hour steps are what the app offers below three hours; above that it offers
whole hours only, so a half hour up there is untested rather than known-bad. A
Number cannot change step part-way, and turning this into a Select of the app's
sixteen values would change the entity's domain on every unit that already has
one, so the step stays 0.5 throughout and the comment says why.

The descriptor's own comment used to admit the upper bound was a guess ("only 0
has been observed on hardware"). The guess of 12 was right; the unit it was
expressed in was not.
2026-08-05 12:21:10 +02:00
perseus177 9ee0329467 feat(airconditioner): reset the legacy filter counter locally
FilterCleanAlarm_Clear, through the same single-token options merge as every
other setting on /mode/vs/0. Measured on an ARTIK051_KRAC_18K: 2.04 Changed and
FilterTime_95 (9 h 30 min) -> FilterTime_0, still zero on a fresh DTLS session
and on every poll after; none of the other 17 tokens moved and the alarm
entries stayed Deleted.

The counter has had no reset until now, and the descriptor said so: two earlier
rounds against live hardware failed, and the conclusion drawn from them was
that the reset had to be cloud-only. That conclusion was wrong, and the way it
was reached is the interesting part -- it came from diffing every resource the
appliance reports before and after pressing reset in Samsung's app, which
showed only the counter zeroing and the alarm clearing. A trigger token cannot
show up in such a diff, because a trigger is never stored. The appliance's own
app sends this token and skips the write when the counter is already zero.

Both failures stay in the comment, because they say what this is not: writing
FilterTime_0 (the value is not writable -- 5595 -> 5595 after 69 s, 1925 ->
1925 after 65 s, two units, opposite power states), and POSTing the cloud
capability's command name to /actions/vs/0 (real name, wrong transport).

Gated on the FilterTime_ token, so it appears only where there is a counter to
reset; newer boards report filter usage through their own resource and would
need a different mechanism.
2026-08-05 01:46:29 +00:00
Marc Billow ccdfe7088e Merge pull request #281 from atc722/agent/nv9000d-regression-fix
Fix read-only sensor categories and add NV9000D coverage
2026-08-04 21:40:03 -04:00
Marc Billow b59b5ae1b4 Merge pull request #290 from perseus177/ac-autoclean-stop
feat(airconditioner): stop a running auto clean, and read its progress
2026-08-04 20:53:44 -04:00
perseus177 290a348017 feat(airconditioner): stop a running auto clean, and read its progress
Three tokens describe the drying cycle these boards run after cooling, and the
switch only covered the first. AutocleanProgress_ is how far a running cycle
has got, and StopAutoClean_ is a channel for ending one early -- its presence
is what says the appliance accepts that at all, which is how the appliance's
own app gates its stop button. Both tokens are reported by the ARTIK051_KRAC_18K
that issue #136 was about, and by every KRAC fixture here.

The percentage scale is the app's own: it renders the token into a
`<progress max="100">` with a "{{value}}%" label beside it. An idle unit reports
1 rather than 0 -- the same floor the laundry firmware's progressPercentage sits
at when Ready -- so 0-vs-1 is not a reliable "is it running" test, and the
button is deliberately not gated on it.

The sensor shares AUTO_CLEAN's catalog entry the way auto_clean_legacy already
shares the switch's: same figure, different board generation, distinct key so
nothing collides if a board ever reported both.

Stacked on #289 (this branch is cut from it) -- rebase or merge that first.
2026-08-05 00:45:20 +02:00
kkqq9320 15279066b5 review: move the state_class into the shared tuple's fourth column
The frozenset was a parallel structure for a per-row fact, and the comment
above the tuple already described it as a fourth column -- so the comment
promised the right shape and the code did something else. Fixed to the shape
the comment described: _AIR_QUALITY_SENSORS carries state_class per row and
the comprehension unpacks it, with _RECORDED_AIR_QUALITY and its duplicated
rationale block deleted.

air_monitor imports the same rows and now unpacks four, but discards the
fourth. That board (issue #210) has stamped all five readings as
`measurement` since it was added; consuming the column would silently drop
long-term statistics for Odor and CleanLevel on shipped devices, which is a
behaviour change this branch has no evidence to make. The grade/concentration
split stays scoped to the air purifier.

test_shared_sensor_tuple_keeps_its_three_column_shape asserted the premise
this replaces -- that widening the tuple breaks air_monitor's import -- so it
is replaced rather than renumbered: one test that the rows carry their own
state_class, and one that air_monitor still imports and still stamps all five.
2026-08-04 15:25:12 +09:00
hoon 3d0dca20f9 Add NV9000D cooktop coverage and fix sensor setup 2026-08-04 14:47:42 +09:00
kkqq9320 96d06369bc review: floor the sensing interval at one minute
Dropping native_min to 0 fixed the read range and quietly opened a write:
native_min governs what the user can enter, not just what renders, so 0
became enterable and would have gone out as periodicSensingInterval "0".
Nothing establishes what that does to this board -- both fixtures report 600,
the app's smallest choice is 10 min, and 60 s is the lowest value confirmed
accepted. The two precedents leaned on differ in exactly the way that matters:
oven.cook_time and operational.delay_start_hours sit at a zero floor under a
value where 0 is a real setting ("no timer", "no delay").

One minute is also the resolution this board reports results at.
lastSensingTime lands on an exact minute on every sample from the AVT-WW-TP1
and A-VTWW-TP2 boards -- both fixtures, plus eleven consecutive live readings
-- where the TP1X/AC/hood boards report arbitrary seconds. A sub-minute
interval is unobservable here whether or not the board honours it.

So native_min goes to 1 rather than 0, and the read rounds up instead of to
nearest so a sub-minute reading renders as 1 rather than falling below the
entity's own floor. The write still refuses anything under a minute -- a None
return, the silent no-op range_hood._lamp_level_write uses for a level the
device didn't advertise -- since native_min only guards the UI path, not a
service call.
2026-08-04 14:44:54 +09:00
kkqq9320 a5484f746a review: unfold AI Purify into one entity per field
Review feedback on #268. The largest change is that the sensing-mode select
no longer folds two device fields into one control.

periodicSensingActivationState and autoExeState are independent knobs, and the
appliance presents them that way -- its own UI has an on/off for AI Purify
separately from the three mode choices. Folding them lost two things: a
configured action was invisible while the feature was off, and no select
option could toggle the feature without also overwriting the action. The
switch was not the duplicate it looked like.

So the switch now owns periodicSensingActivationState alone, and the select
owns autoExeState alone. That resolves the hardcoded-options finding at the
source rather than working around it: the select reads supportedAutoExeState
via options_field -- the same shape SOUND_MODE already uses for
supportedModes -- instead of carrying a typed-in tuple, so a board advertising
a fourth action is accepted on both the options list and the write path.
_sensing_mode, _sensing_mode_write and _SENSING_MODE_BODIES are all gone with
the fold.

Option slugs are now the advertised values lowercased (off / airpurify /
alarm) rather than invented names. The catalog carries the labels, so the two
'off's stay distinguishable in the UI: the switch's means the unit isn't
sampling, the select's means it samples and doesn't act on the reading -- what
the app calls "sensing only".

Also from the review:

  * _interval_minutes checks `is None` so a reported 0 stays 0, and native_min
    drops to 0 since sub-30s values round there. oven.cook_time and
    operational's delay hours are the precedent -- both convert a device time
    value and floor at zero. The Number-rather-than-Select choice is now
    stated in the write helper: the app offers three fixed intervals, but this
    resource advertises no supported-values or range field (supportedAutoExeState
    sits right beside it, so the board does advertise constraints where it has
    them) and it accepted 60 s, six times finer than the app's smallest choice.
  * _skip_time_write no longer splices a malformed half back onto the wire.
    The read side already refuses one it can't parse; the write side now
    zeroes it to match.
  * air_sensing_state and last_air_sensing_level lose enabled_default=False,
    matching range_hood.AIR_LEVEL_CHECK. Hiding two of three read-only keys
    while claiming key parity with that capability -- and leaving the third
    visible -- had no justification behind it.
  * The catalog-parity test drops periodic_air_sensing from its key set: that
    key is a SwitchDesc here and a BinarySensorDesc on the hood, so the two
    live in different platform catalogs and are worded differently. The claim
    now covers only the three read-only sensor keys, where it holds.
  * Tests route through the descriptors (_desc(key).write_fn / .value_fn)
    rather than module-private helpers, matching test_air_monitor_capabilities.

startSensingOnce stays unbound, now explicitly rather than by omission -- the
module comment records it as deferred. It looks like a one-shot "sense now"
button, but this board acknowledges writes it discards, and nothing has
confirmed the side effect yet.

Goldens are untouched: the key set is unchanged, only sensing_mode's value
moves from the folded slug to the raw autoExeState.
2026-08-04 14:20:58 +09:00
kkqq9320 412fff9b99 feat(air_purifier): expose the AI Purify sensing engine on /airlevelcheck/vs/0
/airlevelcheck/vs/0 has been covered as "periodic air-quality sensing
scheduler plumbing" since the registry gained a coverage stub for it. Two
AVT-WW-TP1-23-AXX500 dumps (issues #84 and #190) show it is not plumbing: it
drives the feature the SmartThings app calls AI Purify, where the unit wakes
on a timer, samples the air, and optionally acts on the result. Every field
is named, none are opaque, and two of them are already user-set on the
reported units.

The select's three on-states are the app's own options rather than an
invented grouping -- it offers exactly "Sensing only" (sample, take no
action), "Auto clean" (purify while the air reads bad, stop once it
improves) and "Get notified" (raise a SmartThings notification). Labels were
transcribed from the Korean app and rendered in English; the auto-stop half
of "Auto clean" is the app's own description and is not otherwise visible in
the dump, which reports only the selected autoExeState. The remaining entity
names follow their raw fields rather than inventing a concept -- the skip
window is "sensing skip", after periodicSensingSkipStatus/Time.

Three of this registry's four board families report the resource with the
same field names -- TP1X_DA-AC-AIR (#130), A-VTWW-TP2 (#151) and AVT-WW-TP1
(#84, #190). Only ARTIK051_TVTL (#56) has no such href, and its golden is
unchanged. Bound unconditionally rather than behind a match_fn; the one field
that genuinely varies (periodicSensingInterval, absent on the #130 board) is
gated per-entity, so that board gets eight entities instead of nine rather
than a broken one.

range_hood.AIR_LEVEL_CHECK already models this same href, and its read-only
keys are reused verbatim here so both families share one catalog entry. It is
deliberately not imported: the hood exposes periodic_air_sensing as a
read-only BinarySensorDesc and this board needs a writable SwitchDesc on that
key, so reusing the hood's capability would migrate every hood user's entity
to a different platform.

Every write was exercised on AVT-WW-TP1-23-AXX500 hardware. This board hands
out 2.04 for writes it silently discards (see HEPA_FILTER's filter-reset
note), so an echo proves nothing -- each was judged by whether the value
survived a reconnect, which forces a new DTLS session, fresh discovery and a
fresh observe of the href, leaving no cached state to read back:

  * sensing_mode's combined two-field PUT lands both fields, both ways:
    sensing_only -> auto_purify raises autoExeState with activation still On,
    and back again lowers it.
  * The sensing-skip switch holds Off -> On and back.
  * The half-preserving time writes hold: from 13:00-23:00, writing
    start=07:30 then end=22:00 left the device on '07302200' -- each write
    kept the half it wasn't given.
  * periodic_air_sensing and sensing_interval: writing 60 s drove an observed
    ~60 s sensing cycle.
  * The read side of the skip window is separately cross-confirmed on two
    units: #84's sits at the inert '00000000', #190's carries a real
    '03002300' (03:00-23:00), which is what pins the HHMMHHMM split.
  * The other two families get the writes on field-shape grounds -- the same
    basis on which they already share MODE, HEPA_FILTER and the air-quality
    sensors.

range_hood._timestamp moves to common.epoch_to_utc so both callers share it,
matching how filter_usage_percent was shared. No behaviour change.

Every existing entity is untouched: the three golden updates are purely
additive, no renames, no unit or device_class changes.
2026-08-04 12:57:03 +09:00
Marc Billow b5699badbe Stop the v1 migration re-keying devices onto a placeholder serial
_serial_from_unique_id took the entry's unique_id at face value. That is
right for an entry whose unique_id holds a real serial, but the unique_id
records what the config flow believed when it ran, not what the registry
holds now -- and for two firmware families those are different things.

Entries added before the placeholder rules landed (issues #83/#189) were
keyed on the placeholder itself: `localthings_Nothing(SVC)` for the
ARTIK051_DONGLE_REF dongles, `localthings_FFFFFFFFFFFFFFF` for the
DA_WM_A51_20_COMMON laundry boards. The coordinator has been resolving
those same boards to the host ever since, so their devices and entities
are host-keyed today. Migration read the placeholder back off the
unique_id, decided the host-keyed rows were the stale ones, and rewrote
them onto the placeholder -- reintroducing exactly the collision those
issues exist to prevent, since every unit of the family reports the same
placeholder and would go back to sharing entity unique_ids.

Run the recovered string through resolve_serial, which is the whole point
of that helper being shared. The old `host:port` special case stays: it's
a config-flow-history artifact rather than a device-reported serial, so
resolve_serial can't recognize it.

The repair pass had a second, narrower way to lose data. Removing a device
takes its entities with it (entity_registry.async_device_modified), and
the removal branch ran after the entity pass -- so an entity that had just
been re-keyed rather than removed, because its serial-keyed key was free,
was destroyed a few lines later along with the entity_id, name and area
the rewrite existed to preserve. Move surviving entities onto the device
they now belong to before removing the duplicate.

Reachable when the serial-keyed device exists but a given entity's
serial-keyed key doesn't -- e.g. the user deleted the visible duplicate by
hand, which is the first thing anyone hitting #236 tries.

Also fold the modelNum `<model>|<board>` split into resolve_model beside
resolve_serial. The config flow and _run_discovery each had their own copy
under a comment promising they matched; a device that renames itself on
the first poll is what a drift there looks like.
2026-08-04 00:12:40 +00:00
Marc Billow 6033709f24 Replace the blanket "cannot connect" with a real failure taxonomy
Adding a device had one message for nearly every way it could fail: "Cannot
connect to the device. Verify the IP address is reachable and the CA
credentials are correct." That covers an IP with nothing on it, an
appliance on cloud-only firmware, a device still holding the session from
the last attempt, a device that answered and rejected our certificate, and
Home Assistant having no internet to reach Samsung's cloud. Only one of
those is fixed by checking the IP and the CA credentials, and the message
gave no way to tell which one you had.

The probe already gathers enough to tell them apart, so classify it:

- cert_rejected     the appliance sent a certificate alert. The CA
                    credentials aren't the AC14K_M CA it trusts, or they
                    don't pair. Far and away the most common real setup
                    mistake, and previously indistinguishable from a typo
                    in the IP address.
- handshake_failed  a fatal alert unrelated to the certificate (protocol or
                    cipher mismatch) -- no amount of fiddling with CA
                    credentials will fix it.
- handshake_timeout the ClientHello probe proved a DTLS server is right
                    there, but the handshake never finished. Usually the
                    appliance is still holding the association from a
                    previous attempt; it clears on its own in about a
                    minute.
- ports_closed      ICMP port-unreachable on the whole range: something is
                    at that address and it isn't exposing a local API.
                    Cloud-only firmware (TCP 8888 only) lands here.
- no_dtls_server    some ports open|filtered, none speaking DTLS -- likely
                    another device on that IP.
- no_response       nothing came back at all.
- cloud_unreachable couldn't reach Samsung's cloud gateway for the UUID.
                    An internet problem on HA's side, not the appliance's.
- unexpected_response  authenticated fine, then returned something we can't
                    read. Neither connectivity nor credentials.

Certificate alerts are read back out of the error text OpenSSL puts in
DtlsCoapSession's ConnectionError, not by re-probing. The library's
diagnostic probe would report the alert authoritatively, but it drives the
handshake far enough to commit association state on the device -- and an
orphaned association is exactly what makes the *next* attempt time out
(RFC 6347 4.2.8), which is a bad trade on a path the user is about to
retry.

Telling ports_closed from no_response needs the UDP sweep to separate a
refusal from an unreachable. Both leave a port "not live", but ECONNREFUSED
is a *response* -- the host is there -- while EHOSTUNREACH/ENETUNREACH mean
the datagram never left. A wrong IP on the local subnet never answers ARP
and fails every send that way, so treating the two alike would have told
those users their appliance was on cloud-only firmware. The sweep now
returns live/refused/unreachable separately, and the preferred-port rescue
moved out of it into _sweep_ports: the rescue is a candidate-selection
decision, and folding it into the sweep's verdict destroyed the evidence
the message is built from.

Every failure carries the error key that fits it, so the flow maps
exceptions instead of guessing, and logs the specifics (alert name, per-port
outcome, response code) at warning level -- the messages that mention the
log now have something to point at.

Certificate re-minting for a reused leaf is also narrower and more correct
as a result: it now triggers on CertRejected specifically, rather than on
"every attempt raised ConnectionError and a port was confirmed".

All five translation catalogs carry the eight new messages. The non-English
ones are my own work rather than a native speaker's; corrections welcome.
2026-08-03 20:29:53 +00:00
Marc Billow 15be379243 Rebuild device discovery on the ClientHello probe and resolve identity up front
Two problems, one setup path.

Port detection (issue #211): the config flow found the DTLS port by
elimination -- a 1-byte UDP probe can't tell a silent port from a real
DTLS server, so every port it couldn't rule out got a full certificate
handshake, and every false positive cost the whole 12s HANDSHAKE_TIMEOUT_S
before the next was tried. Adding an appliance took 30-40s.

smartthings-local 0.1.2 ships a stateless ClientHello probe that settles
this positively: a real DTLS server answers with a HelloVerifyRequest in
~1 RTT, and per RFC 6347 4.2.1 it does so without allocating association
state, so the probe leaves nothing behind on the appliance. The whole
49152-49160 range is probed at once and exactly one confirmed port is
given a certificate handshake. Fanning out is safe here in a way racing
real handshakes is not -- each probe is bounded by a 3s budget, so the
pool costs one probe's wall clock rather than the sum of the range, with
no losing threads left running behind us.

The UDP sweep stays as the fallback for when the probe confirms nothing:
it errs in the opposite direction (it reports everything it can't rule
out), so it still surfaces a device on a path that eats our ClientHello,
and it keeps its issue #192 preferred-port rescue.

Port detection now runs first and needs no credentials, so an unreachable
host fails before any round trip to Samsung's cloud. And a second
appliance reuses the existing entry's leaf cert rather than re-minting --
every device accepts the same one -- which makes adding one independent
of Samsung-cloud reachability. A confirmed-live device rejecting the
reused leaf (the UUID does rotate) re-mints and retries once, so reuse
stays self-correcting; a timeout doesn't, since a fresh cert can't fix
nothing answering.

Identity (issue #236): the coordinator seeded device_serial with the
configured host and only replaced it after the first successful poll. But
device_serial mints *permanent* registry keys -- entity unique_ids and
device identifiers -- so anything registering before that poll returned
was written into the registry keyed on the IP address forever. The
connection-mode sensor is added unconditionally rather than from `bound`,
so it was the reliable victim: when the serial-keyed identity appeared
moments later HA created a second device and entity, and the IP-keyed
pair was orphaned. Deleting them didn't help; the next restart that lost
the race recreated them.

The probe already learns the identity, so store it on the config entry --
serial, model, manufacturer, device type. The coordinator seeds
device_serial and its DeviceInfo from those at construction, so keys are
correct from the first entity that registers even if the first poll is
slow or fails outright. There is no placeholder left to correct.

Discovery now treats the registered identity as authoritative rather than
re-keying a device that already has registry entries; it adopts and
persists the polled identity only for an entry that has none, and warns
if a different appliance answers on the same IP.

Entry version 1 -> 2 recovers the serial from the entry's unique_id (the
flow has always keyed it on the probe's serial) and repairs what the old
registration orphaned: IP-keyed devices and entities are rewritten in
place where the serial-keyed key is free -- keeping entity_id, name, area
and every automation referencing them -- and removed where both exist,
since the IP-keyed one has been dead since the restart that made it.
Placeholder-serial boards (issues #83/#189) were keyed two ways at once,
`host:port` on the entry and `host` in the registry; migration collapses
the entry onto the registry's form. One resolve_serial() now serves both
sides, so they can't drift apart again.

The remaining step in the desired pipeline -- probe for subdevices, then
register devices, then populate entities -- already holds:
_enumerate_subdevices_blocking runs before _run_discovery, which runs
before platforms are forwarded. Duplicating it in the config flow would
mean re-running Pattern B's per-href fallback probe, which is the
opposite of what issue #211 is about.
2026-08-03 20:14:54 +00:00
Marc Billow 6a6eef25b8 Fix ty type error in test_dryer_drum_clean.py
for_device_by_model returns DeviceRegistry | None; accessing .capabilities
directly off the inline call result left the None case unnarrowed. Switched
to the same reg/resources-tuple helper pattern every other by-model test
file in this suite already uses, which ty resolves cleanly.
2026-08-03 19:38:57 +00:00
Marc Billow da25d567cb Remove pointless catalog-literal tests; document the anti-pattern
Two tests added while triaging #244/#226 just re-asserted a translation
string against the catalog entry that had been written moments earlier
(dryer_cycle_table_03's '51'/'53'/'4e', dishwasher_cycle's '83'/'86').
Neither exercises any code path -- they pass by construction and only
break when someone later edits the label text for wording, not when the
actual code/value mapping regresses. tests/test_translations.py already
holds the invariants that matter for catalog data.

Documents the anti-pattern in the adding-device-support skill so future
translation-only fixes don't reach for this pattern again.
2026-08-03 19:32:44 +00:00
Marc Billow e89aa4bab5 Fix transposed Normal/Express 60 dishwasher cycle labels (issue #226)
'83' and '86' were swapped in the dishwasher_cycle catalog. Both the
original DW9000F-class fixture this table was built from and the issue
#226 reporter's board share the identical DeviceType_0812, and the
original fixture's own editCourseList puts the two codes back to back
(positions 4-5) -- a plausible adjacent-pair transcription slip. The
reporter's live confirmation (selecting 'Normal' ran the physical Express
60 program and vice versa) settles which way: '86' is Express 60, '83' is
Normal.

The energy-sensor part of the same issue was already resolved per the
issue thread (the device genuinely doesn't report usage, so the sensor's
removal was correct) -- not touched here.
2026-08-03 19:30:02 +00:00