Two Samsung air purifiers of the same model report the identical, well
formed serialNum `BS7SP9AW400114A` (issue #381). Since the entry's
unique_id, the device registry identifiers and every entity unique_id
were all minted from that string, the second unit was refused as already
configured, and would have collided entity-for-entity even if it hadn't
been.
This is the third firmware family to ship an unusable serialNum, after
`Nothing(SVC)` (#83) and the flash-unset sentinel (#189), and the first
one no heuristic can catch: the value is well formed, it's just shared.
`is_placeholder_serial` was a dead end.
So identity moves onto /oic/d's `di`, falling back to /oic/p's `pi`, then
the serial, then the host. `di` is what the protocol already uses to
address the endpoint -- if it were wrong or shared, OCF discovery and the
DTLS association wouldn't work at all -- and it's device-scoped, where
`pi` is platform-scoped and would be shared by a board hosting several
logical devices. Both units in #381 report a distinct `di`. A board that
answers neither resource lands exactly where it did before, so no
existing hardware regresses.
The re-key can't happen in async_migrate_entry: the UUID is only readable
from the device, and an entry can load entirely from its snapshot while
the appliance is off (#295). So v3 -> v4 only records the legacy key, and
the coordinator adopts the UUID on the first live poll, rewriting the
entity registry, the device registry (including subdevice identifiers)
and the entry's unique_id together. Rewriting rather than recreating is
what lets a user keep entity_ids, names, areas, statistics and every
automation that references them.
Three rules keep that adoption from misfiring:
- A poll that reads no UUID never demotes a UUID-keyed entry back onto
its serial, so one failed reconnect doesn't re-key every entity.
- A changed UUID is followed only when the serial still corroborates it
(a factory reset may regenerate `di`) or when the entry was keyed on
its IP, which was never an identity to defend.
- When the identity is rejected as a different appliance, the serial
isn't adopted either -- otherwise the intruder would gain exactly the
corroboration needed to win the next poll.
Also stop redacting `di`/`pi` from diagnostics. They're randomly assigned
per-unit UUIDs, not account data, and blanking them is what made the
first #381 diagnostics download unable to answer the only question it was
requested to answer. The owner-set device name stays redacted.
Fixes#381
- _on_cloud_courses_changed cleared the canonical-view cache but never
pushed state: select.py's current_option reads coordinator.data,
which only moves on async_set_updated_data, so a change here (the
new toggle, or apply_cloud_courses naming a program -- which has
called this same method since before the toggle existed) sat stale
in the UI until an unrelated poll or observe happened to run next.
Now calls _push_cache_snapshot() too.
- _refresh_cloud_course_issue now runs from __init__.py's
options-update listener on every entry save, not just a
cloud-course-specific one. Saving an unrelated option
(CONF_BYPASS_REMOTE_CONTROL, say) before this device's first poll,
or while it's rehydrated offline, read /course/vs/0 as empty --
indistinguishable from "nothing pending" -- and would delete a
Repair a real poll had every reason to raise. Now a no-op on an
empty rep, leaving whatever issue state already exists untouched
until a real poll can judge it.
- The "cloud_courses" menu's off-state note was a raw English string
built in config_flow.py and substituted via description_placeholders
into all 7 locales' descriptions -- unlike the SmartThings screen
names quoted elsewhere (deliberately English everywhere; that's a
third-party app's own label, not ours), this one named LocalThings'
own "Offer downloaded cycles"/"Device settings" labels, which are
translated per locale and should have matched. Replaced with a
permanent, state-independent sentence translated in the catalog
itself, in all 7 locales, instead of conditional Python-built text.
Also restores a word an earlier edit dropped from
async_step_cloud_courses's docstring ("can complete confidently").
Tests: two new regression tests, each confirmed to fail against the
pre-fix code before being fixed -- one drives coordinator.data through
a toggle via a fixture already sitting on a one-time cloud override, so
current_option actually depends on the cloud store instead of falling
back to the raw course code; the other simulates a second coordinator
against the same entry with an empty resource cache (a not-yet-polled
restart) and confirms an existing Repair survives an unrelated option
save. Full suite (1585 tests), ruff, and `ty check custom_components
tests` all pass.
Issue #364: several reporters got the "downloaded cycles not set up"
Repair despite never meaning to use the feature -- one device appears
to auto-populate a slot from a SmartThings-provided example. The two
reporters who did complete setup successfully both hit the same root
cause for their earlier failures: the SmartThings app has two
similarly-named screens ("Cycle", which lists everything including
local courses, and "Download cycles", the one that actually matters
here), and nothing in our instructions said to use the second one
specifically.
Global disable (CONF_CLOUD_COURSES_ENABLED, entry.options, default
on):
- New coordinator.cloud_courses_enabled property, mirroring
CONF_LEARN_MODES' shape -- off stops the Repair and stops offering
already-named programs as cycles, without discarding anything
already learned or named.
- Deliberately does NOT stop _observe_cloud_courses' passive recording:
guided/manual setup depend on live observation to detect a newly
selected program at all, and leaving it running means turning the
option back on immediately surfaces anything set up in the meantime
instead of asking the user to redo it. Documented on the const and
on the property.
- New __init__.py options-update listener calls a new
coordinator._on_cloud_courses_changed(), which both clears the
canonical-view cache (memoized, so a stale view would otherwise keep
answering with pre-toggle state -- caught by two failing tests
before this) and refreshes the Repair. Nothing else needed this
because every other option is read live on its own next use; cloud
courses is the only one with standing Repair/cache state to refresh
immediately rather than on the next unrelated change.
- Toggle exposed in Device settings as "Offer downloaded cycles",
alongside prose explaining why some devices show the Repair
unprompted.
Instructions, in every shipped locale (en/de/es/it/cs/nl/ko) --
otherwise a locale missing the new/changed strings would silently show
English or the old text, the same gap issue #376 already tests for:
- Every guided-setup screen, the manual edit form, and the Repair
itself now say explicitly: open the SmartThings app (not the
appliance), and tap "Download cycles" specifically -- a separate row
from "Cycle", further down the screen -- not the general cycle
picker. Also states plainly that the appliance doesn't need to be
nearby or running the cycle, just powered on and connected.
SmartThings' own screen names are kept in English in every locale
(verified only in the English app via the reporter's screenshots;
translating them without evidence of what Samsung's own localized
app shows would be a guess this codebase's translations otherwise
avoid).
- The Repair's description now also points at the new toggle for
anyone who doesn't want the feature at all.
- The "cloud_courses" menu screen shows a note when the option is
currently off, since guided/manual setup still work in that state
but nothing named there will appear as a selectable cycle until it's
turned back on.
Tests: coordinator-level tests cover the option defaulting on,
suppressing a new Repair, clearing an already-open one, hiding/
restoring the cycle-select entry as the option flips (which caught the
canonical-cache bug above), and that passive observation keeps running
regardless of the option. Options-flow tests cover the new field's
default and that it persists. Translation catalog tests
(test_every_language_mirrors_the_english_catalog et al.) cover every
locale's topology and placeholders for the changed/added strings.
Full suite (1583 tests), ruff, and `ty check custom_components tests`
(CI's exact invocation) all pass.
HA logs a removal warning (2027.8) every time this name is accessed on
releases that carry UnitOfDensity, attributed straight to this
integration since it's a plain module-level import. UnitOfDensity is
the replacement, but hacs.json's floor (2025.1.0) predates it existing
at all -- pytest-homeassistant-custom-component 0.13.316, the newest
available, still has no UnitOfDensity either, so this can't be a
static import on either branch without breaking support for part of
the version range.
Resolved with a runtime getattr instead: reads UnitOfDensity off the
homeassistant.const module if present and uses its
MICROGRAMS_PER_CUBIC_METER member, otherwise falls back to the plain
(un-deprecated, on those older releases) constant. The getattr
short-circuits before the deprecated name is ever touched on a release
new enough to have UnitOfDensity, so the warning stops firing there
without dropping support for anything still on the old one. Same
feature-detection shape _relabel_particulate_statistics already uses a
few lines down for new_unit_class.
Verified the resolution logic directly: against the installed HA
(2026.2.3, pre-UnitOfDensity) it resolves to the plain constant with no
warning; a simulated future homeassistant.const with UnitOfDensity
present resolves to it without ever touching the deprecated name (a
guard that raises on that access never fires).
Full suite (1573 tests), ruff, and `ty check custom_components tests`
(CI's exact invocation) all pass.
06's Korean text ('이불') is identical to the confirmed Bedding codes
24/6f, and Bedding reads better than the guessed 'XXL Laundry' wording
issue #342 originally gave it. Applied across all 7 locale catalogs
(matching each locale's own already-translated Bedding text, not a
fresh translation) and folded into issue #376's WF21T6500KV test as a
21st confirmed code instead of a flagged exclusion.
test_confirmed_washer_table_02_missing_course_names (#342) updated to
match; its docstring now notes 06's wording was later corrected by
#376 rather than pinning the old value as if still current.
Issue #376 reported Korean UI labels for 21 washer (Table_02) and 18
dryer (Table_03) codes that had no translation and were rendering as
raw hex in the UI, from a WF21T6500KV washer (DA_WM_A51_20_COMMON) and
DV19T8745BV dryer (DA_WM_TP1_21_COMMON).
Cross-checked every reported code against translations/ko.json before
translating anything: several share their exact Korean text with a
code the catalog already has a confirmed label for (washer '19'/'AI
맞춤세탁' matches '2b'/'69'; dryer '3a'/'살균건조' matches '21'; dryer
'3c'/'피트니스' even matches washer '2f', a cross-table reuse; etc.) --
those reuse the established label instead of a fresh translation. The
rest (Wool/Lingerie, Boil Wash, Soft Bubble, Padding Care, and others
with no catalog precedent) are new translations of the reporter's
Korean text.
One code is deliberately NOT applied: washer '06' ('이불', Bedding per
this report) conflicts with 'XXL Laundry', already locked in by
test_confirmed_washer_table_02_missing_course_names (issue #342). Two
reports of the same nominal Table_02 disagreeing on one code is a real
discrepancy, not a wording question -- left alone pending the reporter
(or another Table_02 owner) confirming which device's '06' is actually
wrong, same caution as the existing '24'/'33' transposition history
(issue #343).
Added to all 7 locale catalogs (en/de/es/it/cs/nl/ko), not just
English: HA falls back to English for any key a locale is missing, so
translations/en.json alone would still pass
test_every_language_mirrors_the_english_catalog's topology check while
leaving every other locale showing English text for these codes.
Tests: two new tests lock in the English labels and, for every reused
code, that every locale's label actually matches its anchor code (not
just English) -- the same gap issue #343 fell through, since the
topology test alone can't catch a locale-specific mistranslation.
Full suite (1575 tests), ruff, and ty all pass.
Two releases landed since the 0.1.6 pin, both confirmed by upstream
(QuiteYellow, in issue #361) as additive/opt-in with no interface
changes on our side:
- 0.1.7: server-certificate profiles (SamsungServerProfile), a bounded
DTLS handshake deadline (connect() now defaults to a 12s bound
instead of none), and a cancellable connect() via
ConnectCancellation. Our connect() call sites in coordinator.py and
config_flow.py pass no args, so they pick up the bounded handshake
for free; the cert-profile and cancellation pieces are opt-in and
unused here.
- 0.1.8: fixes blockwise OBSERVE notification reassembly
(QuiteYellow/SmartThings-Local#39) -- a notification carrying only
the first Block2 block was previously handed straight to
on_notification instead of being reassembled, and separately, the
Block2 loop could append a retransmitted/late block as if it were
the next one, or miscompute the next block offset after a mid-
transfer size downshift. Both corrupt a multi-block observed
resource without necessarily truncating it -- the "premature end of
stream" / "error decoding unicode string" CBOR failures reported in
issue #361 on /mode/vs/0. All error types stay within the existing
compatible-built-in table (ConnectionError/TimeoutError subclasses),
so no exception handling changes.
`>=0.1.6` already permitted pip to resolve 0.1.8 on a fresh install,
but an environment that already has 0.1.6 or 0.1.7 satisfying that
floor won't be upgraded by Home Assistant's requirement check -- which
is what #361's reporter is very likely still hitting on 0.22.0.
Raising the floor to >=0.1.8 forces that upgrade on the next release.
Verified against smartthings-local 0.1.8 from PyPI: full suite (1573
tests), ruff, and ty all pass. No source changes needed beyond the
three version pins (manifest.json, requirements-dev.txt, Dockerfile).
The OutdoorTemp_ options token was only surfaced on legacy boards
(is_legacy_board), even though 14 of 17 fixtures carrying the token are
non-legacy. issue #367 confirmed with a 48h field capture (r=0.92
against weather.forecast_home) that the token tracks real outdoor
temperature independent of board generation, and that no non-legacy
board exposes an alternative outdoor-temperature resource.
Split a token-presence-only exists_fn (_has_option_token_any_board) for
this token, leaving _has_option_token's legacy gate untouched for the
other options[] settings that still need it. The -55 offset itself was
only field-validated on Celsius-locale boards, so a second gate
(_reports_celsius, reading the board's own /temperatures/vs/0) keeps
the sensor off the one Fahrenheit-locale fixture on record rather than
guess whether the same offset and unit still apply there. Ships
enabled_default=False since multi-split installs report the same token
on every indoor head, which would otherwise create one duplicate active
sensor per head.
Updates the golden fixtures for the 13 affected Celsius-locale boards
and the artik051_krac test that had asserted outdoor_temperature stays
off newer boards; adds coverage for the Fahrenheit-locale gate.
The OutdoorTemp_ options token was only surfaced on legacy boards
(is_legacy_board), even though 14 of 17 fixtures carrying the token are
non-legacy. issue #367 confirmed with a 48h field capture (r=0.92
against weather.forecast_home) that the token tracks real outdoor
temperature independent of board generation, and that no non-legacy
board exposes an alternative outdoor-temperature resource.
Split a token-presence-only exists_fn (_has_option_token_any_board) for
this token, leaving _has_option_token's legacy gate untouched for the
other options[] settings that still need it. Ships enabled_default=False
since multi-split installs report the same token on every indoor head,
which would otherwise create one duplicate active sensor per head.
Updates the golden fixtures for the 14 affected boards and the
artik051_krac test that had asserted outdoor_temperature stays off
newer boards.
Widen async_rehydrate's guard to cover the identity and Subdevice rebuild,
not just the replay. A stored row missing a field the current dataclass
declares raised KeyError straight out of async_setup_entry, which only
handles ConfigEntryNotReady -- so the entry landed in SETUP_ERROR, which HA
never retries, with its DTLS session left open on the fixed source port the
next attempt binds. It now fails the same way an unreachable device does.
Write the snapshot immediately instead of through async_delay_save. A
deferred write outlives whatever queued it: removing an entry inside the
delay window deleted the file and then had it recreated, orphaned, when the
timer fired; and a reload scheduled by _reconcile_rehydrated read the
pre-reload snapshot back off disk, so a device going quiet again mid-reload
rehydrated the stale set and reconciled a second time. Banking it before the
reconcile fixes the ordering. Failures are logged rather than raised -- a
board reporting something the JSON encoder rejects must not break polling.
A coverage gap is a claim about what the device currently reports, so
replaying a discovery snapshot shouldn't make it. Offline it would restate
the last live poll's conclusion while pointing the user at a diagnostics
download that stays empty until the appliance answers, and any drift in the
resolved device name between snapshot and live would churn the issue.
Not deduplication: HA already keys issues on (domain, issue_id), preserves
dismissed_version across async_get_or_create, and reloads non-persistent
issues with their dismissal intact -- one row per entry, and an "Ignore"
survives restarts.
An appliance switched off at the wall used to take its whole config entry
down with it: async_setup_entry raised ConfigEntryNotReady, so the device
read as failed and its entities existed only as registry rows until the
appliance came back.
Loading the entry anyway isn't enough on its own. Entities here are the
output of discovery, discovery only runs inside a successful poll, and
platforms enumerate `bound` exactly once at forward time -- so an entry
that loads while offline loads empty, and with no listeners subscribed the
base coordinator stops rescheduling and never polls again.
Bank the resources dict each successful first cycle hands _run_discovery,
along with the subdevice candidate list and the /oic identity that route
the registry, and replay it through _run_discovery when the first refresh
fails. Storing the poll input rather than a rendered entity list keeps one
implementation of discovery instead of two: the offline entity set is
produced by the same code that produced the live one.
Three things fall out of that:
- Platforms judge entity existence against `discovery_resources`, not the
live cache. The live cache deliberately stays empty, which is what keeps
a restored entity `unavailable` rather than rendering a stale value for
an appliance nobody can currently reach.
- A live discovery that disagrees with the snapshot reloads the entry --
platforms can't adopt a changed set in place, so a firmware update or a
sibling subdevice that starts answering needs a fresh setup.
- The entry holds one coordinator listener for its lifetime, so polling is
scheduled regardless of how many entities are live.
An entry that has never reached the device has no snapshot, keeps raising
ConfigEntryNotReady, and closes its session on the way out as before -- no
metadata to build a device from, and it leaves room for setup flows that
need to interact with the appliance (#168).
Restores the two tests PR #303 rewrote, narrowed to that no-snapshot path.
PR #303 loads the entry when the first poll fails. Measured on its
branch, that produces an entry with zero bound entities and zero
coordinator listeners, so DataUpdateCoordinator never reschedules and
the device never recovers without a manual reload.
Record why entities can't be created offline here (discovery is the only
source of `bound`, and platforms enumerate it once), what a working
version would need (persisted discovery snapshot, reconcile-on-reconnect,
a listener that keeps polling alive), and the cheaper retry-and-reload
option that solves the filed issue on its own.
When a device is offline or unreachable during Home Assistant startup,
previously raised . HA's built-in
retry mechanism uses exponential backoff up to 15 minutes, which leads to
a poor user experience for local LAN devices.
Catch initial connection errors during and log a
warning instead of failing setup. This allows platforms to set up and
entities to be created (in an unavailable state), while the coordinator
continues background retry polling.
The investigation is written around the FilterTime_<N> option token on
/mode/vs/0, and its conclusion holds for the ARTIK051_KRAC_18K it was
measured on. An ARTIK051_PRAC_20K has no such token: no FilterTime, no
FilterAlarmTime, no FilterCleanAlarm anywhere in its options blob. It
keeps the counter in /filter/airdustfilter/vs/0 as a percentage of a
500-hour interval instead.
Neither route resets it. FilterCleanAlarm_Clear to /mode/vs/0 returns
4.00 with the options blob byte-identical; writing filterUsage as the
string "0" returns 4.00; writing it as an integer returns 5.00. That
last difference is the useful part — two payloads differing only in JSON
type returning different codes rules out an unresolved href or an
unrecognised field name, leaving read-only as the reading.
Adds a scope line to the intro and a section documenting the board, its
two resource dumps, the attempt table, and an observation-only
workaround for percentage-counter boards.
Confirmed on a live dishwasher's sensor.*_progress history (not in any
shipped fixture): Prewash -> Wash -> Rinse -> Drying ->
DryingWithDoorOpen -> Finish -> Idle. The door-open drying-assist stage
had no catalog entry, so it fell through to the sensor.py fallback and
rendered as the raw lowercased string 'dryingwithdooropen' instead of a
readable name.
No code change needed -- operational.py's progress SensorDesc already
derives its options from the catalog via translated_states(), so a new
state key is picked up automatically.
#366 added Dutch state translations for dryer_type ('Electricity') and
progress, but predates #371's lowercase-key normalization: its progress
states were already superseded (every one it added is already in nl.json's
current 'state' table, lowercase), and its dryer_type addition used the
pre-#371 capitalized key.
dryer_type itself was never wired for translation at all -- no
device_class, no options -- so nothing in any locale's dryer_type.state
table was ever read. Give it device_class=enum, options=("electricity",)
(the only value confirmed across shipped fixtures), and the same
value_fn=lower() normalization #371 used elsewhere, then add the
'electricity' state across all seven locale catalogs, not just Dutch.
Home Assistant resolves a catalog `unit_of_measurement` against the default
language, not the user's (entity_platform re-fetches 'en' for exactly this
key), because a unit is part of the state's identity -- the recorder writes
it into statistics metadata and compares it across restarts. Localizing it
would make switching Home Assistant's language look like a unit change and
suppress the sensor's statistics.
So the six non-English entries were never read, and the English one only
restated what `unit="cycles"` already said. Same displayed unit either way;
this drops seven catalog keys that looked translatable but weren't.
Translates status values the integration previously surfaced as raw
Samsung strings: cycle progress ('Rinse' -> "Rinsing"), diagnosis state,
buzzer volume options, and the drum-clean counter's unit. Progress and
diagnosis become enum sensors so Home Assistant looks their state up in
the catalog; progress keys come from lowercasing the device's own value
rather than a hardcoded map, so adding a language is a catalog-only
change (PR #341 review).
Rebased onto main, which has since gained the issue #345 sticky hold, and
fixed up for two problems that combination exposes:
Home Assistant refuses an enum state that isn't in the sensor's options,
which takes the entity out rather than degrading it. The sticky hold froze
progress at the device's raw 'Finish' while rep_fn had been normalized to
'finish', so every completed cycle -- the exact path #345 exists to serve
-- would have produced a rejected state.
Separately, options built from the catalog can only ever list values we
have a translation for, while this registry's rule is that an unrecognized
device value renders raw. Every progress token the shipped fixtures
advertise is covered today, but Samsung ships more devices than we have
dumps for, so the sensor platform now admits the live value into its own
options: known values translate, unknown ones display untranslated instead
of breaking the entity.
The drum-clean unit moves from a native unit to the catalog because Home
Assistant rejects an entity declaring both. Note it resolves against the
default language, so the localized unit strings are inert -- kept only
because every catalog must mirror English key for key.
Also moves the diagnosis normalizer to common.py, so dryer.py doesn't
import a private symbol from dishwasher.py to get it.
Review follow-up on the v2 -> v3 migration.
`new_unit_class` only exists from HA 2025.11, but hacs.json still declares
2025.1 as the minimum. On anything in between, naming that keyword is a
TypeError raised out of async_migrate_entry, which fails the config entry
outright -- the integration would not load at all for those users. The
keyword is now feature-detected, and the relabel is wrapped so that no
recorder-side surprise can cost anyone the integration: it is a
convenience, and without it they simply get Home Assistant's own
units_changed repair, which is where they were before this existed.
The version bump also no longer happens when the recorder wasn't loaded.
That case isn't distinguishable from an install without the recorder, but
burning the one-shot migration on a boot where it merely failed to come up
would leave the statistics suppressed permanently, so the entry stays on
v2 and the next start retries.
The instance-suffix regex was wrong: discovery.instance_suffix yields
`_<n>`, so keys are `dust_1`, not `dust1`. Unreachable today because
AIR_QUALITY binds an exact href, but the comment claimed a guarantee the
pattern didn't provide and the test pinned a form that can't occur.
Also drops a stale claim in airconditioner.py that air_monitor rejects the
pm10/pm25/pm1 mapping, which is no longer true as of this branch.
Tests cover both signatures, the deferral and its retry, and a relabel
that raises. The older-HA guard is mutation-checked: removing the feature
detection fails it.
Adding a unit to a sensor that recorded long-term statistics without one
is not cosmetic: Home Assistant raises a units_changed repair and then
*suppresses statistics generation* for that entity until a human resolves
it (sensor/recorder.py's compile path hits `continue`). Shipping the PM
device classes on their own would therefore have silently frozen the very
history the labels were meant to describe.
A v2 -> v3 config-entry migration relabels the statistics metadata first.
It rewrites only the metadata row, never the recorded values -- these
readings were always µg/m³ and only the label was missing, so nothing
needs converting, which is why this uses async_update_statistics_metadata
and not change_statistics_unit. It reads entity_ids back off the entity
registry rather than rebuilding them from descriptor keys, since a renamed
entity's statistic_id no longer follows from its key, and it is scoped by
device family: range_hood and airconditioner still declare no unit for
their identically-named sensors, so relabelling theirs would create the
exact mismatch this exists to prevent.
With the migration in place there is no longer a reason to hold the labels
back on air_monitor, so it takes them too. That board is one of the two
whose fixtures pin the grade bands the mapping rests on -- typing the
purifier and not the monitor was an inconsistency, not caution. It still
declines the shared state_class column, unchanged.
Recorder coupling is guarded: after_dependencies pulls it in when
configured, the import is local to the migration, and a setup without it
is a no-op. Freshly created entries mint v3 directly, having no history to
relabel.
Tested at both levels. The unit-level tests patch the recorder and assert
which entities are relabelled with which arguments, covering the renamed
entity, subdevice-prefixed and instanced keys, the near-miss keys
(dustbag_/dustbin_), the skipped families and the recorder-absent path.
Because a mock can only prove the call is made, not that it does what the
migration needs, a second suite drives a real in-memory recorder end to
end: statistics seeded unitless, migration run, metadata confirmed to read
µg/m³ / concentration, recorded means confirmed byte-identical, and the
units_changed issue confirmed present before and absent after.
Review follow-up to PR #365, which landed the washer 0A/B0 labels and the
air-purifier PM device classes.
The three particulate units were spelled with U+00B5 MICRO SIGN. Home
Assistant's DEVICE_CLASS_UNITS holds only the U+03BC GREEK SMALL LETTER MU
spelling, so every purifier logged a per-entity "not a valid unit for the
device class" warning asking the user to file a bug against us. The two
characters render identically, and the PR's own test hardcoded the wrong
one, so the test agreed with the bug. That test now takes the expected
unit from HA's own constant, and a new registry-wide guard
(test_sensor_device_class_units.py, mirroring the SwitchDesc guard from
issue #349) checks every SensorDesc unit against HA -- these were the only
three invalid pairs among 29.
Getting this in before release matters more than usual: the recorder
writes unit_of_measurement into long-term statistics, so correcting it
afterwards would raise a "units changed" repair for anyone who had run the
released version.
Also settles what the second element of a dust reading's value[] is, which
was the open question behind issue #325's request for another dump. It is
the device's own graded air-quality level: it appears only on the fields
carrying a magnitude (Dust/FineDust/SuperFineDust/CO2) and not on
Odor/CleanLevel, which are grades already; it reads 0-2 against index 0's
0-31; and CleanLevel equals the highest per-field grade on 9 of the 11
fixtures reporting the resource. It stays unbound -- ARTIK051_TVTL grades
good air as 0 while every other family uses 1, so a shared descriptor
would need a per-family offset -- but it is what confirms the PM mapping
without relying on field names: 18 grades one step above the floor as
SuperFineDust yet sits at the floor as Dust, on two families that both
floor at 1, so the firmware itself treats the three fields as different
scales ordered coarse-to-fine. Each field's floor boundary also brackets
the Korean CAI band for its tier (PM10 at 30/31, PM2.5 at 15/16). Pinned
against the shipped fixtures in test_air_quality_grade_column.py.
air_monitor keeps its untyped sensors, but the docstring now gives the
real reason: the evidence carries over, and what is deliberately deferred
is the statistics migration for entities shipped unitless since issue #210.
Smaller fixes: en.json's "Mixed load" -> "Mixed Load" to match the
catalog's title casing and issue #363's own wording; de/ko gave B0 the
same string as the existing "34" Mixed, so a machine exposing both showed
two identical options; washer.py's shared-label list still said "'24'
Towels", which went stale when issue #343 found 24/33 transposed; and
0A/B0 now have a locale-wide translation guard like every other confirmed
code batch.
Table_02 codes 0A (Towels) and B0 (Mixed load) were reported for a
WW90DG5G34ABLE (issue #363). Air-purifier Dust/FineDust/SuperFineDust
map to PM10/PM2.5/PM1 in µg/m³ from a same-moment SmartThings
correlation (issue #325).
* vacuum_station: bind VS9700 stick battery via /status/stick/vs/0
* vacuum_station: translate stick labels; drop diagnostic category
Address review: localize stick entity names in non-English files, and
keep wand status/BLE as primary entities rather than diagnostic.
---------
Co-authored-by: Nicolas <11050206+WiestDaessle@users.noreply.github.com>
An independent review of the last two commits' diff turned up five real
problems and one CI-breaking one. Fixed all of them:
- tests/test_coordinator_error_handling.py assigned directly onto
coordinator instance attributes (coordinator._poll_once = dict), which
`ty check custom_components tests` -- what CI actually runs, not the
narrower `ty check custom_components` this branch had only been
spot-checked against -- flags as invalid-assignment. Switched to
monkeypatch.setattr, matching every other new test on this branch.
- coordinator.py's subdevice-enumeration failure comment claimed the
probe "retries naturally next cycle." It doesn't: _run_discovery sets
self._discovered = True unconditionally later in the same cycle, which
is what gates the whole block, so a failure here is a first-and-only
attempt, not a retried one -- a composite appliance's sibling
subdevices are missing for the config entry's lifetime until reload.
(A separate flag to retry wouldn't actually fix that either: every
platform's async_setup_entry enumerates coordinator.bound exactly
once, so a later-successful enumeration still couldn't add entities
without a reload.) Corrected the comment and raised debug to warning,
since the effect is silent and permanent otherwise.
- config_flow.py's new diagnostic-handshake fallback (_resolve_alert)
was being called once per failing candidate, inside _handshake_and_read's
scan loop -- contradicting _diagnostic_alert's own docstring ("only
runs once every real candidate has already failed"). Two real costs:
up to CLIENTHELLO_PROBE_TIMEOUT_S extra latency per failing port on
the sweep-fallback path (several candidates), and -- more seriously --
the diagnostic commits DTLS association state on each port it touches,
which can make _probe_and_validate's own CertRejected re-mint retry
(a fresh _handshake_and_read call against that same scan) time out
against the very port it just polluted, per the RFC 6347 §4.2.8
concern already documented elsewhere in this file. Moved to a new
_diagnose_failures helper called once, after the loop, against the
single best (confirmed-live) candidate.
- _resolve_alert also ignored the alert's level: ProbeResult.alert is
set for a received alert of either severity, but only a fatal one (2)
means the appliance broke off the handshake over it -- a warning
(e.g. close_notify) was being read as a rejection reason. Older
library exception text never had this ambiguity (OpenSSL only renders
an exception for a fatal alert), so this was a bug the redaction
fallback introduced. Now filters to level == 2.
- async_raw_read's new HomeAssistantError wrapper reported a translated
error but left a confirmed-dead session installed, so every
subsequent read/write would keep failing identically for up to the
next full poll interval. Now closes the session on any non-TimeoutError
failure, matching _poll_once's own posture.
- async_raw_write_sequence's verify_after fallback (vcode, vrep = 0, {})
is indistinguishable from a real 4.04 by raw_code alone. Added a
read_error field so a caller can tell "couldn't verify" from "the
device said no."
Tests: 8 new/extended (once-not-per-port diagnostic count, alert-level
filtering, the warning log + permanent-loss framing, session-closing on
both async_raw_read and the verify_after path, read_error surfacing).
Full suite (1527 tests), ruff, ruff format, and -- critically --
`ty check custom_components tests` (the CI-matching invocation) all pass.
Excessive for user-facing docs -- this is implementation detail
(_defer_reconnect_for's tolerance logic) that belongs in the
coordinator's own comments, where it still lives, not in "Known
device behavior". No functional change.
Answers a review callout on the 0.1.6 upgrade that wasn't actually
addressed, only mentioned in a commit message: a dead reader now raises
SessionClosedError (a ConnectionError) instead of hanging a request out
to its timeout and surfacing as an ambiguous TimeoutError. Mechanically
this was already routed correctly -- SessionClosedError isn't a
TimeoutError, so _defer_reconnect_for's isinstance check already skips
its multi-cycle tolerance for it -- but nothing recorded *why*, and the
callout's whole point was that downstream (i.e. this repo) should hear
about the resulting timing change explicitly, not infer it from the
dependency bump.
Before 0.1.6, a truly dead reader was indistinguishable here from a
slow blockwise transfer: both could only ever surface as a TimeoutError,
so _POLL_TIMEOUT_LIMIT's multi-cycle tolerance (~2 minutes at the
default 30s interval) was the only thing standing between a genuinely
dead session and a reconnect. Now that the library confirms reader
death directly, that failure mode skips the tolerance and reconnects on
the very first occurrence -- intended, and strictly faster recovery,
but a real change in observed timing worth calling out for anyone
correlating reconnect-log cadence with device behavior.
- _poll_once, _defer_reconnect_for, and the _POLL_TIMEOUT_LIMIT comment
now say so directly, cross-referencing each other.
- README's "Known device behavior" section gets a paragraph so this
isn't only visible to someone reading the coordinator's source.
- Two new unit tests pin the distinction directly:
_defer_reconnect_for(SessionClosedError()) is False (no tolerance),
while SessionTimeoutError keeps the existing _POLL_TIMEOUT_LIMIT
tolerance -- so a future change can't quietly merge the two paths
back together.
Full suite (1523 tests), ruff, and ty pass.
A follow-up review of the 0.1.6 upgrade found four call sites where a
library exception (new typed one or the old bare ConnectionError/
TimeoutError it replaced) could escape this integration's own
reconnect/logging or a service call's translation layer entirely,
instead of being handled the way equivalent failures already are
elsewhere in this file:
- _attempt_observe_mode's own _connect_session() reconnect (fires only
when the session was closed out from under it concurrently) had no
try/except, and neither did either of its two call sites in
_async_update_data. A failure there escaped uncaught: HA's
DataUpdateCoordinator has its own final safety net so nothing crashed
the config entry, but non-TimeoutError failures logged a full ERROR
traceback instead of this integration's deliberately quiet "poll
failed, reconnecting" voice, and skipped its own reconnect bookkeeping
entirely. Fixed by catching around just the connect call (the only
unguarded raise path in the method -- subscribe_hrefs/
await_observe_notifies already handle their own failures), landing in
the same "give up on push this cycle" state abandon_observe_attempt()
already produces for the subscribe-failed and stale-session branches.
Deliberately does not touch _close_session() (self._session is
already None here -- _connect_session only ever publishes it after a
full success), _reconnect_is_frequent() (that window records the poll
path's own reconnects; feeding it a secondary path's failure would
over-trigger its warning threshold), or _resubscribe_due (that flag
means "a live session nothing has tried yet" -- setting it here would
re-enter the doomed handshake every cycle instead of letting
_last_observe_attempt_ts pace the retry).
- _enumerate_subdevices_blocking's _connect_session() call (first
discovery only) had the same gap. Fixed the same way: log and fall
through on the resources _poll_once already returned this cycle,
rather than losing first discovery over a failed subdevice probe.
- async_raw_read (backing the read_resource service) had no exception
handling at all -- a session/network failure during a live debug read
reached the service caller as a raw, untranslated library exception,
unlike write_resource's equivalent path. Now wrapped the same way
async_send_command/async_raw_write_sequence already are, raising
HomeAssistantError with a new debug_read_failed translation key
(added to all 7 shipped locales).
- async_raw_write_sequence's verify_after tail sat outside the method's
own try/except, so a failed confirmation read discarded the write
results that had already landed by throwing past them. Now caught
per-href inside the verify loop instead: a failed read is treated the
same as a 4.04/empty one (held=None, "couldn't verify" -- not lost or
misreported as a revert), and the rest of the batch still gets
checked.
Design for the first fix (the trickiest -- it's mid-lock, and has to
interact correctly with observe-mode state and the poll path's own
bookkeeping without corrupting either) was worked through with a
dedicated review pass before implementing.
Tests: new coverage for all four (test_coordinator.py's
test_attempt_observe_mode_survives_a_failed_reconnect, a new
test_coordinator_error_handling.py for the subdevice-enumeration case,
and two additions to test_services.py for the read-service and
verify_after cases). Full suite (1521 tests), ruff, and ty all pass.
Bumps the smartthings-local floor from >=0.1.2 to >=0.1.6 (manifest,
requirements-dev.txt, Dockerfile) and adopts the interface/behavior
changes introduced along the way:
- 0.1.3 ("redacted typed failures", PR #23) replaced connect()'s
ConnectionError(f"DTLS handshake error: {e}") with fixed, redacted
exceptions (SessionError, SessionTimeoutError, etc.) that never carry
backend text -- including the TLS alert name. The config flow's
_classify_handshake_failure relied on parsing that text out of the
exception (_alert_name) to tell a rejected certificate from any other
handshake failure; against a current library that regex never matches
again, silently downgrading every setup failure to the generic
"cannot_connect" message.
Fixed by adding _resolve_alert: it still tries _alert_name first (a
harmless fallback if it ever matches), then falls back to one bounded
smartthings_local.protocol.dtls_probe.diagnose_dtls_handshake() call
against the specific port that failed, which classifies the fatal
Alert straight from the raw record instead of an exception string.
_handshake_and_read now threads the resolved per-port alerts into
_classify_handshake_failure, so CertRejected vs. HandshakeFailed keeps
working the way it did before the redaction.
- 0.1.3 also moved the DTLS session onto a connected UDP socket (see
endpoint.py's open_connected_udp_socket), which changes why
coordinator._local_source_port needs a unique port per device -- the
kernel now demuxes by the full local-port/remote-peer tuple instead of
relying on an unconnected recvfrom(). Docstring updated to match.
- 0.1.6's reader-thread fail-fast fix (_check_live/_reader_running) makes
a dead reader raise SessionClosedError immediately instead of hanging
a request out to its timeout. No code change needed: SessionClosedError
is a ConnectionError subclass (not TimeoutError), so the coordinator's
existing _defer_reconnect_for/isinstance(e, TimeoutError) split already
routes it to the immediate-reconnect path.
Every other new/changed piece (endpoint.py, dtls_probe.py bounded
probing, auth.py's CertificateAuth/PskAuth providers) stays behind
compatible built-in exception types and unchanged get()/post()/
subscribe()/ping() signatures, per the library's own compatibility
table, so the coordinator's and observe.py's broad exception handling
needed no changes.
Tests: added coverage for _resolve_alert's exception-text vs.
diagnostic-handshake fallback, _classify_handshake_failure with a
resolved alerts mapping, and an end-to-end config-flow re-mint test
against a FakeSession that raises the new redacted SessionError instead
of the old text-bearing ConnectionError.
Full suite (1517 tests), ruff, and ty all pass against smartthings-local
0.1.6 installed from PyPI.
queue_get accepts dict | list since #335's Collection test started
queueing a batch list, but _get_reps was still typed list[dict] --
ty flagged the list.append() as invalid. No behavior change.
`_raw_read_blocking` decoded the CBOR body and kept it only when it was a
Property map, so a Collection -- which answers the `[devcol rep, {href,
rep}, ...]` batch `parse_device0_batch` reads -- came back as `2.05` with
`rep: {}`. That renders as "the resource exists and has nothing in it",
which is the opposite of what a populated batch means, and `/device/0`
itself would have read the same way.
It cost a real result: issue #335's board answers `/sec/devices` (the
`x.com.samsung.devcol` sibling of `/device/0`, and the one remaining place
a composite appliance could be enumerating its indoor units) with exactly
that empty-looking 2.05, and it was nearly written off as a dead end.
The read path now returns the decoded body alongside `rep`, and the service
response carries it as `body` whenever it isn't the map `rep` already has --
omitted for the ordinary case rather than duplicating every rep in every
response. Records the probe round this came out of: indexed leaves 4.04 on
that board, and the UUID prefix confirmed routable by a positive control, so
Patterns A/B/C are ruled out there on evidence rather than on absence.
Issue #335's board reports a sibling in subdeviceIdList and then 4.04s on
all 26 seeds enumerate_subdevices tries, which reads like "there is nothing
there". Comparing every captured /oic/res in the corpus says otherwise: only
the ARTIK051_DONGLE_FAC_18K board advertises its operational tree at all.
The other five list the onboarding surface and stop -- the range board hides
a live /device/1 behind a ten-link /oic/res -- so an href's absence from
/oic/res is not evidence, and Pattern A's index scan is dead weight
everywhere except the board it was written against.
What that leaves untried is the bare indexed leaf: every indexed href this
project has ever seen arrived inside a /device/<n> batch, and /device/1
4.04ing is evidence about the Collection, not about /mode/vs/1. Leaves
without their Collection is already confirmed BORA behavior in the other
namespace (issue #205). Records the probe list, the OCF composite-device
clause that suggests /sec/devices, a positive control for whether the UUID
prefix routes at all, and the dead ends worth not re-treading.
Adds washer_cycle_table_00 and dryer_cycle_table_00 translation catalog
entries, confirmed by the issue #357 reporter selecting each cycle on a
WF45R6300AW/US washer and DVE45R6300W/A3 dryer and reading back the raw
course code. Table_00 is a separate, older course-code family from the
existing Table_02/Table_03 catalogs -- laundry.cycle_select's table_href
scoping already keeps them apart, so this is a translations-only change.
Table_00 was previously used only as an example of an unconfirmed table in
tests; those now use Table_99 for that role, and new tests assert the
confirmed codes translate and that the resolved key routes to the new
table-scoped catalog entries.
Mirrored to all shipped languages (cs/de/es/it/ko/nl) to keep
tests/test_translations.py's key-for-key invariant.
The reporter's machine_state history for the same two cycles flips to
idle on the exact second progress reads 'Drying' (12:08:51 and 14:01:26),
which settles what the previous commits had to infer: rep_fn returns
'Idle' whenever state isn't active, so it cannot have produced that
value, and the only remaining path was the ungated sticky_live_fn the
bypass returned in its place. Replaying the sequence against the pre-fix
path reproduces the reported Cooling, Finish, Drying, Idle exactly; the
fix holds Finish through it.
Comments only -- swap the inference for the observation that confirms it.