A switched-off washer or dryer fails in the DTLS handshake, not in a
poll: `_poll_once` opens the session itself, so `_connect_session` runs
to its 12s timeout with nothing to show. The poll path treated that like
any other poll failure and ran its reconnect -- close the session, pause,
poll again -- but there is no session to close and no association for the
device to clean up, so the retry was the identical handshake five seconds
later. That cost 29s of every 30s interval, and the same again on every
setup attempt for an entry with no snapshot to load from.
`_poll_once` now records which of the two failed, and the poll path skips
the retry when the handshake is what never completed. A session that
opened and then broke still reconnects within the cycle.
The log was the half the reporters saw: an ERROR every cycle (plus a
WARNING once three "reconnects" piled up) for a state this integration is
built to sit through, which issue #269's reporter read as the integration
having failed. An outage now reports once, DEBUG for the cycles after it,
and INFO when the device answers again.
Fixes#269
Widen async_rehydrate's guard to cover the identity and Subdevice rebuild,
not just the replay. A stored row missing a field the current dataclass
declares raised KeyError straight out of async_setup_entry, which only
handles ConfigEntryNotReady -- so the entry landed in SETUP_ERROR, which HA
never retries, with its DTLS session left open on the fixed source port the
next attempt binds. It now fails the same way an unreachable device does.
Write the snapshot immediately instead of through async_delay_save. A
deferred write outlives whatever queued it: removing an entry inside the
delay window deleted the file and then had it recreated, orphaned, when the
timer fired; and a reload scheduled by _reconcile_rehydrated read the
pre-reload snapshot back off disk, so a device going quiet again mid-reload
rehydrated the stale set and reconciled a second time. Banking it before the
reconcile fixes the ordering. Failures are logged rather than raised -- a
board reporting something the JSON encoder rejects must not break polling.
A coverage gap is a claim about what the device currently reports, so
replaying a discovery snapshot shouldn't make it. Offline it would restate
the last live poll's conclusion while pointing the user at a diagnostics
download that stays empty until the appliance answers, and any drift in the
resolved device name between snapshot and live would churn the issue.
Not deduplication: HA already keys issues on (domain, issue_id), preserves
dismissed_version across async_get_or_create, and reloads non-persistent
issues with their dismissal intact -- one row per entry, and an "Ignore"
survives restarts.
An appliance switched off at the wall used to take its whole config entry
down with it: async_setup_entry raised ConfigEntryNotReady, so the device
read as failed and its entities existed only as registry rows until the
appliance came back.
Loading the entry anyway isn't enough on its own. Entities here are the
output of discovery, discovery only runs inside a successful poll, and
platforms enumerate `bound` exactly once at forward time -- so an entry
that loads while offline loads empty, and with no listeners subscribed the
base coordinator stops rescheduling and never polls again.
Bank the resources dict each successful first cycle hands _run_discovery,
along with the subdevice candidate list and the /oic identity that route
the registry, and replay it through _run_discovery when the first refresh
fails. Storing the poll input rather than a rendered entity list keeps one
implementation of discovery instead of two: the offline entity set is
produced by the same code that produced the live one.
Three things fall out of that:
- Platforms judge entity existence against `discovery_resources`, not the
live cache. The live cache deliberately stays empty, which is what keeps
a restored entity `unavailable` rather than rendering a stale value for
an appliance nobody can currently reach.
- A live discovery that disagrees with the snapshot reloads the entry --
platforms can't adopt a changed set in place, so a firmware update or a
sibling subdevice that starts answering needs a fresh setup.
- The entry holds one coordinator listener for its lifetime, so polling is
scheduled regardless of how many entities are live.
An entry that has never reached the device has no snapshot, keeps raising
ConfigEntryNotReady, and closes its session on the way out as before -- no
metadata to build a device from, and it leaves room for setup flows that
need to interact with the appliance (#168).
Restores the two tests PR #303 rewrote, narrowed to that no-snapshot path.
PR #303 loads the entry when the first poll fails. Measured on its
branch, that produces an entry with zero bound entities and zero
coordinator listeners, so DataUpdateCoordinator never reschedules and
the device never recovers without a manual reload.
Record why entities can't be created offline here (discovery is the only
source of `bound`, and platforms enumerate it once), what a working
version would need (persisted discovery snapshot, reconcile-on-reconnect,
a listener that keeps polling alive), and the cheaper retry-and-reload
option that solves the filed issue on its own.
The investigation is written around the FilterTime_<N> option token on
/mode/vs/0, and its conclusion holds for the ARTIK051_KRAC_18K it was
measured on. An ARTIK051_PRAC_20K has no such token: no FilterTime, no
FilterAlarmTime, no FilterCleanAlarm anywhere in its options blob. It
keeps the counter in /filter/airdustfilter/vs/0 as a percentage of a
500-hour interval instead.
Neither route resets it. FilterCleanAlarm_Clear to /mode/vs/0 returns
4.00 with the options blob byte-identical; writing filterUsage as the
string "0" returns 4.00; writing it as an integer returns 5.00. That
last difference is the useful part — two payloads differing only in JSON
type returning different codes rules out an unresolved href or an
unrecognised field name, leaving read-only as the reading.
Adds a scope line to the intro and a section documenting the board, its
two resource dumps, the attempt table, and an observation-only
workaround for percentage-counter boards.
`_raw_read_blocking` decoded the CBOR body and kept it only when it was a
Property map, so a Collection -- which answers the `[devcol rep, {href,
rep}, ...]` batch `parse_device0_batch` reads -- came back as `2.05` with
`rep: {}`. That renders as "the resource exists and has nothing in it",
which is the opposite of what a populated batch means, and `/device/0`
itself would have read the same way.
It cost a real result: issue #335's board answers `/sec/devices` (the
`x.com.samsung.devcol` sibling of `/device/0`, and the one remaining place
a composite appliance could be enumerating its indoor units) with exactly
that empty-looking 2.05, and it was nearly written off as a dead end.
The read path now returns the decoded body alongside `rep`, and the service
response carries it as `body` whenever it isn't the map `rep` already has --
omitted for the ordinary case rather than duplicating every rep in every
response. Records the probe round this came out of: indexed leaves 4.04 on
that board, and the UUID prefix confirmed routable by a positive control, so
Patterns A/B/C are ruled out there on evidence rather than on absence.
Issue #335's board reports a sibling in subdeviceIdList and then 4.04s on
all 26 seeds enumerate_subdevices tries, which reads like "there is nothing
there". Comparing every captured /oic/res in the corpus says otherwise: only
the ARTIK051_DONGLE_FAC_18K board advertises its operational tree at all.
The other five list the onboarding surface and stop -- the range board hides
a live /device/1 behind a ten-link /oic/res -- so an href's absence from
/oic/res is not evidence, and Pattern A's index scan is dead weight
everywhere except the board it was written against.
What that leaves untried is the bare indexed leaf: every indexed href this
project has ever seen arrived inside a /device/<n> batch, and /device/1
4.04ing is evidence about the Collection, not about /mode/vs/1. Leaves
without their Collection is already confirmed BORA behavior in the other
namespace (issue #205). Records the probe list, the OCF composite-device
clause that suggests /sec/devices, a positive control for whether the UUID
prefix routes at all, and the dead ends worth not re-treading.
The Download-course candidate came from "Course_ read while a non-sentinel
one-time payload is loaded". On the first rep after any restart that is
indistinguishable from a payload left over from a previous run, so an
appliance holding cloud payloads while sitting on an ordinary course
proposed that ordinary course as the Download one. Accepting the prefill
would then make selecting a downloaded program start, say, a cotton wash.
Only a transition actually watched counts now. "Never observed" is a
distinct state from "observed, nothing loaded" -- absent-then-loaded is a
genuine selection and still counts -- so a restored store deliberately
re-enters the unobserved state, since a restart cannot tell the two apart.
Payloads are still learned from that first rep either way; which programs
exist is device fact regardless of when they were loaded. It is only the
inference about which course means Download that needs the timing.
Both corpus dumps taken off the Download course show the appliance clearing
its one-time token to the FFFF sentinel, so this may never fire on these
boards. That is a reason to expect them to behave, not to depend on it.
Its bytes are not payload slots everywhere. On the DW5000C dishwasher all
four (8E 8D 8F 02) are course codes in that device's own course list, three
already translated -- Plastic, Pots and pans, Baby Care. There the token
marks which ordinary courses came from the cloud; they select with a plain
Course_ write and need no payload, which is consistent with it carrying no
payload token at all. It also has a DownloadCourseList_ token the washers
lack. On both washers the slots share zero overlap with the course list and
a payload is required to select one.
So the "this is not washer-only" claim was wrong, and gating on
advertised_slots offered that dishwasher's owner a naming flow for programs
that already work and are already named. The Repairs card was spared only
because the payload gate added earlier happens to catch it.
cloud_slots() subtracts the device's own course list, which separates the two
readings without guessing at families: what remains is slots that cannot be
selected any other way, which is what this module is for. Everything
user-facing now gates on that -- the options menu entry, the naming flow, the
Repairs count. The dishwasher gets nothing, both washers are unchanged.
Found by reading the dishwasher fixture's options array while answering a
question about it, which is also why diagnostics now reports advertised and
cloud slots separately: the difference between them is the whole distinction.
A survey of every laundry diagnostics dump attached to an issue turned up 14
devices, 4 of which carry cloud-course tokens. Two were already known; the
two new ones are both useful, and one contradicts something the
investigation write-up asserted.
A DW5000C dishwasher (issues #113/#123) advertises four downloaded programs
and carries no payload token for any of them. That is a shape the corpus
didn't have: the feature is not washer-only (DA_DW, not DA_WM), and a device
can name programs whose payloads have never been observed. The existing code
already handles it correctly -- nothing learnable, nothing offered, gap still
counted for the Repairs issue -- so this adds the fixture, golden, and tests
that keep it that way.
A second WW5000C (issues #259/#343, firmware _B048) holds the same saved
program as the first one's captured "Towels", and the two payloads differ at
exactly one byte: byte 3, 04 against 06. Everything else -- id, slot, all
four varying tag values, the whole tail -- is identical. So byte 3 is neither
a per-board constant nor a property of the program, and the doc's claim that
it is always 04 on this board was wrong.
That is also the strongest argument yet for learning payloads per device: a
catalog keyed on program id would have shipped one unit's byte 3 to the
other. Nothing changes in the implementation as a result -- it never had a
catalog -- but the reasoning is now backed by evidence rather than caution.
Also recorded: both WW5000C units advertise the byte-identical slot list
despite different firmware, so the program set looks factory- or
region-assigned rather than user-curated; and a sentinel's byte 2 equals the
selected course on one dump but not the other, so it stays unused.
Byte-aligning the WA55A7700AV's 16-byte payload against the WW5000C's
20-byte one: identical header, and the first four tag/value pairs are the
same tags in the same order at the same offsets -- the part that carries
per-program data has one shape on both boards. The whole width difference
is two trailing pairs the WA55 doesn't carry, and on the WW5000C that
trailing section is byte-identical across all nine programs, so it isn't
program data at all.
Doesn't change the conclusion -- the four shared tags carry non-overlapping
value ranges between the boards, so the encoding is still board-specific and
blobs are still replayed whole. Also records why the WA55's /washer/vs/0
readings can't be used to confirm a decode: that unit is on a local course,
not its cloud course.
A washer whose course table includes "Download"/"Downloaded" runs whichever
program the SmartThings cloud last pushed down. Those programs are now
selectable from the ordinary cycle select, so a downloaded Jeans or Sports
cycle can be started without giving the appliance internet access.
The device turns out to enumerate them itself. `CloudExtraCourse_` on
/course/vs/0 lists one byte per downloaded program, and byte 2 of a
program's payload is exactly that slot id -- verified against all nine
programs on the reporter's WW5000C and against the WA55A7700AV dump already
in the corpus. So nothing here is hardcoded: the appliance says which
programs exist, the payloads are learned by watching what it reports, and
the names come from the user.
That last part is unavoidable rather than a shortcut. A payload is only
visible while its program is loaded, and the appliance never reports a name
for one. So cloudcourse.py persists what has been seen (same rationale as
learned.py's mode store), a Repairs issue tells the owner how many programs
are still unaccounted for, and an options-flow step collects the names. A
program appears in the cycle select only once it is both learned and named.
Selecting one issues the only two-token options write in the codebase --
the course token has to switch to Download in the same write, or the
appliance accepts the program token and silently ignores it (confirmed on
hardware). The Download course code is learned by observation but never
applied until the user confirms it: tokens in this array are replaced by
prefix and never evicted, so a stale program token can appear alongside an
unrelated course, and acting on that would start the wrong wash cycle. For
the same reason a stale token is never reported as the running program.
Also of note:
- There is no single "Download" course code. The WW5000C uses 87, the
WA55A7700AV uses 17 -- same Table_02. Any per-table lookup would have
been wrong on one of the only two devices available to check.
- Payloads are replayed byte-for-byte and never decomposed or rebuilt.
Bytes 5/7/9 do decode to temperature/rinse/spin on the WW5000C, 9 for 9,
and produce nonsense on the WA55A7700AV -- so that decode is written up
in docs/investigations/download-cycle.md and not shipped, and the
read-only sensors it would have enabled were dropped.
- The store reaches the registry as a namespaced synthetic field merged
onto /course/vs/0's rep at read time, so exists_fn/rep_fn/options/write_fn
all see it through their existing signatures. It never enters the state
cache, so it can't be polled over, written to the device, or land in a
diagnostics dump.
- A name that would render identically to another cycle in the same
dropdown is rejected in the flow: the select maps a chosen label back to
a raw value by matching display text.
Non-English catalogs carry the new strings in English for now; they need
real translations.
FilterCleanAlarm_Clear, through the same single-token options merge as every
other setting on /mode/vs/0. Measured on an ARTIK051_KRAC_18K: 2.04 Changed and
FilterTime_95 (9 h 30 min) -> FilterTime_0, still zero on a fresh DTLS session
and on every poll after; none of the other 17 tokens moved and the alarm
entries stayed Deleted.
The counter has had no reset until now, and the descriptor said so: two earlier
rounds against live hardware failed, and the conclusion drawn from them was
that the reset had to be cloud-only. That conclusion was wrong, and the way it
was reached is the interesting part -- it came from diffing every resource the
appliance reports before and after pressing reset in Samsung's app, which
showed only the counter zeroing and the alarm clearing. A trigger token cannot
show up in such a diff, because a trigger is never stored. The appliance's own
app sends this token and skips the write when the counter is already zero.
Both failures stay in the comment, because they say what this is not: writing
FilterTime_0 (the value is not writable -- 5595 -> 5595 after 69 s, 1925 ->
1925 after 65 s, two units, opposite power states), and POSTing the cloud
capability's command name to /actions/vs/0 (real name, wrong transport).
Gated on the FilterTime_ token, so it appears only where there is a counter to
reset; newer boards report filter usage through their own resource and would
need a different mechanism.
CONTRIBUTING.md gains a "Code comments" section: comment the why not
the what, keep it to a sentence or two with a pointer to the load-bearing
evidence, don't re-derive a sibling's already-documented reasoning, and
move failed-attempt investigation logs out of inline comments.
Applied that policy across the codebase: condensed sprawling module
docstrings, per-entity essays, and multi-paragraph rationale blocks down
to their load-bearing conclusions, while preserving the actual "why"
(issue numbers, calibration evidence, gotchas, don't-guess rationale).
No functional code changed — verified via diff review, ruff, ty, and the
full pytest suite (1211 passed).
One inline investigation log (the AC filter-reset "tried and failed"
notes) moved to docs/investigations/ac-filter-reset.md rather than being
deleted, per the new guideline on where that kind of record belongs.