With no cycle set the oven reports x.com.samsung.da.desired = 0, and
flatten() published that straight through as target_temp_c. Home
Assistant rejects it against the Number entity's declared 30-270 range
on every publish, which produced 66,899 log errors over three weeks:
Invalid value for number.samsung_oven_setpoint: 0 (range 30.0 - 270.0)
0 is not a 0 degree target, it is the absence of a setpoint, so treat
anything outside the settable band as absent. null lands as unknown on
both the Number and the Setpoint sensor, the way completion_minutes
already reads when the oven is idle. _setpoint applied these bounds on
the write side already; only the read path was missing them.
Adds the first tests for the sample descriptors. One of them pins a
non-obvious asymmetry: the write path snaps to the 5 degree step grid
before bounds-checking, so 29 commits as 30 and 271 as 270, and only 0
is refused outright. The invariant that has to hold is the weaker one,
that every value the write path commits is one flatten() will publish
back, or a write appears to succeed and then reads as unknown.
The refetch path logged only at debug, and the bridge configures logging
at INFO, so a successful re-read and a total failure produced identical
output: nothing. That makes the hardware validation for #39 impossible
to read.
One line per refetch, promoted to INFO when DEBUG_BRIDGE=1 and left at
debug otherwise, naming the href, the one-shot token, the block count,
and the reassembled size. The token is the part that matters: it is what
shows the re-read used a fresh 4-byte token rather than the observation's
1-byte one, which is the assumption the whole design rests on.
Gating on DEBUG_BRIDGE rather than raising the logger keeps the per-block
retransmit lines out of the way, and matches how the module already gates
its frame dump.
A notification carries only the first block of a large representation
(RFC 7959 §2.6). _dispatch_coap handed that block straight to
on_notification, so consumers decoded a truncated CBOR buffer. Reported
twice on /mode/vs/0: #37 and mbillow/localthings#361.
The recovery is a re-read from block 0 on a fresh 4-byte one-shot token,
not a §2.6 continuation. §3.4 rules out reusing the observation's token,
and this server drops a transfer that opens at NUM>0 under a token it
has not seen, so a continuation is the one shape that cannot work here.
A truncated notification is now withheld and queued to a worker thread
that re-reads the resource and delivers the reassembled representation.
When the re-read fails the notification is dropped at debug level and
the poll tiers carry freshness, which is what they already did.
The re-read has to run off the reader thread: _dispatch_coap runs there
and the transfer waits on an event only that same thread can set. The
worker is serialized and paces between transfers, so a notification
storm stays under the firmware request ceiling.
Also in the Block2 loop, now extracted and shared by both paths:
- compare the response's Block2 NUM against the one requested, so a
retransmitted block is no longer concatenated as if it were the next
- compare ETags across blocks (§2.4) and restart once when the
representation changes mid-transfer
- recompute the next block number from the accumulated byte offset when
the server negotiates the block size down
- re-check reader liveness while waiting on a block, so a mid-transfer
reader death fails fast instead of burning the whole timeout
- guard the token counters and _pending with a lock, now that the
session issues concurrent reads of its own
Closes#39
- add dtls_handshake.py (from #34) to the repo-layout tree
- point the issue #16 / #20 notes at SamsungServerProfile / ServerCertificateAuth instead of calling that path unsupported
- list the new certificate-profile, connect-deadline, and session-interruption test modules
The reader loop exited silently on any socket error, leaving conn/sock
set so the session still looked open. Every later get()/post()/ping()
then waited out its full request timeout on a session nobody was
reading, raising SessionTimeoutError on repeat, forever.
This started biting in v0.1.3 (d677c72), which moved to connected UDP
sockets: a connected socket surfaces ICMP errors on recv, so one
ECONNREFUSED from a rebooting appliance now killed the reader.
- Advisory ICMP errnos (ECONNREFUSED/EHOSTUNREACH/...) no longer kill
the reader; the next datagram usually works.
- Real reader exits log at WARNING; close()-driven exits stay quiet.
- A _reader_running Event lets get/post/ping/subscribe/refresh_observes
fail fast via _check_live() with SessionClosedError instead of waiting
out a timeout. Callers that never start a reader are unaffected.
Refs QuiteYellow/SmartThings-Local#37
The distribution checker enforces an exact-contents allowlist; add the
NOTICE file to the wheel dist-info/licenses expectation and the sdist
required set so the packaged trademark notice passes verification.
Add a Trademarks & disclaimer section to the README and a root NOTICE
file stating this is an independent, unofficial project not affiliated
with Samsung, and that Samsung/SmartThings marks are used nominatively.
Ship NOTICE inside the distributed artifacts by adding it to
license-files (wheel .dist-info/licenses/) and the sdist include list.
hatchling 1.31.0 ships only .gitignore in the sdist; older releases also
bundled .hgignore. check_sdist required an exact member set including
.hgignore, so the Validate workflow's package job failed on main once CI
resolved the newer hatchling.
Require the tracked source set plus the fixed metadata files, and accept
.gitignore/.hgignore as optional members either way.
- Note the Fedora/RHEL SHA-1 crypto-policy block in Part 2 and how
setup_cert.py auto-retries / the manual update-crypto-policies remedy.
- Add dtls_probe.py to the repo-layout tree (it was referenced in three
places but missing from the file listing).
- Drop the .venv/ prefix from the Part 1 probe command so it runs against
the pip-installed package before the Part 4 venv exists.
- Expand the tests parenthetical to name the probe, port-resolution, and
cert-signing suites.
* fix(setup_cert): surface openssl errors and work around SHA-1 crypto policy
The signing step forces -sha1 (the AC14K_M chain requires SHA-1-signed
leaves), which Fedora/RHEL's default crypto policy rejects on OpenSSL
3.x. run() also swallowed stderr, so the failure surfaced as an opaque
non-zero-exit traceback with no diagnostic.
- run() now raises CommandError carrying the command and openssl stderr
- mint_cert retries signing with a scoped OPENSSL_CONF enabling
rh-allow-sha1-signatures when the first attempt fails
- main() prints the update-crypto-policies fallback on failure
Fixes#15
* test(setup_cert): cover SHA-1 signing, error surfacing, and crypto-policy retry
Regression tests for the #15 fix:
- full mint_cert flow (SHA-1 leaf, UUID SAN, custom OIDs, chain assembly)
- CommandError surfaces openssl stderr on a genuine signing failure
- signing retries via the SHA-1 override when the plain attempt is blocked
- run() raises CommandError with detail
GitHub is deprecating the Node 20 action runtime; checkout@v4,
setup-python@v5, and upload/download-artifact@v4 all run on it and
were being auto-forced to Node 24 with a warning. Bump each to its
current major (checkout@v7, setup-python@v7, upload-artifact@v7,
download-artifact@v8), all of which run natively on Node 24.
Convert em-dash prose splices to varied punctuation (periods, colons,
semicolons, commas, parens), turn **Label.**-period bullets into
**Label:** colons, drop sentence-spanning bold in "Traps to avoid", and
cut a couple of hollow intensifiers.
No content, facts, tables, code, or links changed (50/50 line diff). Left
as-is: the `## Part N —` headings (heading-anchor stability), everything
inside code/log fences, table N/A cells, and numbered-list
`**Bold** — desc` carve-outs.
A stateless-by-default DTLS ClientHello probe that classifies a host:port
as DEAD/LIVE/COMPLETED/REJECTED in ~1 RTT off the server's first flight,
sitting in front of the full handshake.
Probe (smartthings_local/protocol/dtls_probe.py):
- Stateless liveness mode (default): stops at HelloVerifyRequest and never
sends the cookie'd second ClientHello, so by RFC 6347 §4.2.1 it leaves
no association on the device — safe to run before a real connect.
- Diagnostic mode (stateless=False): drives the handshake further to
capture cipher/cert-chain/CertificateRequest or a fatal Alert, for
OCF-PKI-wall characterization (#16). Kept out of hot reconnect paths.
- Retransmit + retries: services OpenSSL's DTLS retransmit timer so a
single dropped ClientHello no longer reads as a false DEAD.
MQTT bridge (mqtt_demo):
- Stateless pre-flight gate in session_once() rejects a silent/rebooting
device or wrong port in ~3s (retries=1) instead of eating the 12s
HANDSHAKE_TIMEOUT_S per reconnect.
- OCF-band port autodiscovery when OCF_PORT is unset: races the band in
parallel and returns on the first port to answer LIVE (~1 RTT, abandoning
the dead-port probes), cached across reconnects; the stateless gate
leaves no orphan, preserving the fixed-source-port §4.2.8 invariant.
Validated on real hardware (dryer 49155 / oven 49154): parallel discovery
resolves both ports in <1s, connect with no orphan cooldown, and a wrong
pinned port rejected in ~3s.
Tests: probe behaviour (retransmit recovery, stateless single-flight
guard, silent-port flight budget, diagnostic continuation) and bridge
port-resolution (pinned gate, parallel discovery early-exit, cache).
Root-cause fix for the stale-session stall on always-on appliances (#14).
When the client dies without close_notify (crash, SIGKILL), the device
keeps an orphaned DTLS association keyed to the old 5-tuple; a reconnect
from a fresh ephemeral port presents as a brand-new peer, so the orphan
lingers until the device's own timer reaps it (observed 5-15 min).
RFC 6347 §4.2.8 covers exactly this: a ClientHello arriving on an existing
association's 5-tuple means the peer rebooted, and the server must complete
the new handshake and discard the old association. Add an optional
local_port to DtlsCoapSession that binds the UDP source port, and have the
bridge bind base+appliance-index, so every reconnect re-handshakes over the
same 5-tuple and the orphan is evicted instead of waited out.
Bench-verified on live hardware (2026-07-26): RT-OCF accepts the
same-5-tuple rehandshake (oven, dryer: handshake completes over a
crash-orphaned association, reads work immediately). The oven does not
reproduce the fridge stall even with 11 crash-orphaned OBSERVE
registrations, so fridge-side confirmation of the eviction is still needed.
Co-authored-by: vmvarga <garrysuchiy@gmail.com>
- credit @indykoning + note localthings test path for the washer row
- soften DV90T mnid grouping (mnid=0AJT confirmed on DV5000T only)
- de-speculate the same-family note now that a washer is confirmed
- Lead with the pip-installable library; frame the MQTT bridge as a
reference demo. Add a library quick-start (install, DtlsCoapSession
example, in-memory cert_pem/key_pem variant).
- Fix stale protocol/ + ocf/ references to smartthings_local/*; update
the repo-layout tree (nested package, ocf_root_ca.pem, pyproject.toml,
tests/, publish.yml); drop the non-existent auth.py.
- Correct the write-surface trap: reconciliation is a deferred poll, not
a post-write fetch-back (which itself triggered Samsung's revert).
Distinguish hardware-gated parity (power/child-lock/RC-enable) from
the open oven remote-start problem.
- Note the few write surfaces the cloud HA integration doesn't expose
(dryer course, oven setpoint). Drop the achieved collaborators-wanted
callout.
Nest the two library packages under a single import namespace so they
can ship as one distribution:
protocol/ -> smartthings_local/protocol/
ocf/ -> smartthings_local/ocf/ (git mv, history preserved)
- Rewrite all imports protocol.* -> smartthings_local.protocol.*,
ocf.* -> smartthings_local.ocf.* across the ocf modules, mqtt_demo/
(bridge, descriptor, samples), and tests.
- Add pyproject.toml: dist name `smartthings-local`, hatch-vcs versioning
from v* tags, wheel ships only smartthings_local/.
- Add .github/workflows/publish.yml: build + PyPI Trusted Publishing on
v* tags (OIDC, no stored token).
- Force-include protocol/ocf_root_ca.pem via [tool.hatch.build] artifacts:
it is tracked but matches .gitignore's *.pem, so hatchling's VCS file
selection would drop it — and dtls_session.py loads it at runtime.
- Update mqtt_demo Dockerfile COPY and deploy.sh tar allowlist to the
single smartthings_local/ package.
- .gitignore: build artifacts (_version.py, dist/, *.egg-info/).
Validated: pytest tests/ (11 passed), python -m build produces sdist +
wheel with the pem bundled, fresh pip install resolves all nested imports
with the pem readable from site-packages.
- Add ARTIK051_REF_17K row to the tested-combinations table with a
link to aminorjourney's PR.
- New 'Firmware families — a limitation' section under Part 1
explaining that descriptors are firmware-family-specific with no
runtime feature detection, so the wrong descriptor produces
half-broken sensors rather than a clean error.
- New 'Fridge (ARTIK051)' section under Per-appliance notes with the
capability table + firmware-specific observations (port 49155,
minimal /oic/res, vestigial /hass paths, collection-resource door
model vs newer per-instance-resource fridges).
- Update config-keys reference: CLASS list gains 'fridge',
OCF_PORT defaults list gains fridge=49155.
Four failure modes observed in-house since the polling-first refactor
(709fdf4):
1. Half-open DTLS sessions where the socket stays writable but the peer
has gone silent. Ping sends succeed against a wedged peer because
RT-OCF doesn't reliably emit a RST; only successful polls prove the
session is live.
- PollScheduler exposes last_success_ts (bumped on every 2.05).
- KeepaliveTask takes liveness_fn(); ticks fail if no 2.05 in the
last 60s, even when the ping send succeeded.
- Bridge force-closes the session after 120s unreachable so
run_forever() breaks out of sess.join() and reconnects.
2. RT-OCF cascade under load. One wedged path can eat 8s of timeout,
the next tier tick fires immediately and stacks another attempt,
and the device wedges harder.
- PollTier.timeout_s per-tier override (hot=2s, warm=4s, sweep=15s).
- On TimeoutError, the href goes into a 5-60s cooldown via the
existing _defer_until mechanism.
- take_window_stats() now reports successful-poll RTT separately
from a timeout count, exposed as the "Poll Timeouts (window)"
diagnostic entity in HA.
- Active-window throttle: if the previous health window saw >=3
timeouts and is_active=True, drop back to idle cadence -- stops
stacking polls on a stalled responder.
3. OBSERVE table aging across cloud-auth blips. The device stays
DTLS-reachable but the on-device stack clears its observer table
during the blip, so push delivery stays dead even after upstream
recovers.
- New ObserveRefreshTask per bridge; every 6h derregs all current
observer tokens and re-subscribes on the existing session.
4. Oven lamp/door coupling. Oven hardware auto-drives the lamp from
door state but /mode/vs/0 is warm-tier (30s) so HA showed stale lamp
during a cook.
- Track door + lamp value-change timestamps in descriptor_state.
When the door transition is newer, derive lamp from door_open.
When an HA optimistic write is newer, the cache value wins.
Also: ANSI-coloured WARNING/ERROR lines (NO_COLOR=1 opt-out), jittered
reconnect backoff so dryer + oven don't sync up after a router blip.
In-house verification: running on dryer + oven since 2026-06-03.
Closes#2.
Previously the cert minting script lived in local-tools/ (gitignored)
and the README pointed at a cert-only source that didn't include the
private key or upstream chain.
setup_cert.py now lives at the repo root and live-fetches both the
peer UUID (from the relevant TLS server cert subject DN) and the
full AC14K_M + upstream chain bundle (RemoteAccessCA + CECA + ROOTCA)
from a public mirror. Each fetch has an inline workaround if the
network is restricted (UUID=..., AC14K_M_CERT_BUNDLE=...,
BRAYSTORM_URL=...). Modulus-pair check catches a wrong-key mistake
before signing. bootstrap.py removed -- imported a package that was
renamed in commit 709fdf4.
Output files use neutral client.* names. README, .env.example,
docker-compose.yml, deploy.sh, and config.py updated to match.
Provenance receipts in local-tools/cert_provenance.md.
State freshness now comes from a tiered PollScheduler over the persistent
DTLS session; OBSERVE registrations are kept as an opportunistic
acceleration layer. Behaviour is identical online vs air-gapped except
for worst-case freshness latency.
Adds three modules:
- StateCache: single source of truth, source-tagged change events
- PollScheduler: hot/warm/cold + sweep tiers, write-defer past the
fetchback-revert window, per-window RTT/slow-poll tracking
- KeepaliveTask: CoAP empty-CON ping with consecutive-fail detection
driving MQTT availability
Bridge publishes per-appliance diagnostic entities (Push Active, Last
Update Source, Poll Max RTT, Slow Polls, Poll Errors, Stalest Resource
Age, Last OBSERVE Age) under HA's Diagnostic section. Tier cadences
are descriptor-declared, calibrated against measured per-firmware
ceilings (dryer ~14 req/s, oven ~8 req/s via probe_poll_rate_combined.py).
Drops HEARTBEAT_INTERVAL_S in favour of the descriptor-declared sweep
tier; PING_INTERVAL_S now consumed by KeepaliveTask inside the bridge
rather than driven from main.py.
README explains the push/poll split and what happens when the appliance
is blocked from internet.
Major session of local-OCF reverse engineering against the NV7000BS
oven and DV5000T dryer. Surfaces a working set of HA entities for the
oven and resolves several Samsung-quirk regressions in the bridge's
write path.
Key behavioural fixes:
- OBSERVE registrations now use single-byte tokens. Samsung RT-OCF
silently drops registrations with TKL>1; same 4-byte tokens work
fine for GET/POST. Symptom was that writes returned 2.04 but the
appliance never pushed state changes.
- Per-session random starting tokens + MID. Samsung retains observer
state across DTLS reconnects from the same cert; reusing tokens on
reconnect silently no-ops.
- OBSERVE deregister sent on DtlsCoapSession.close(), with a stop-
watcher thread in PushBridge.session_once() so SIGINT/SIGTERM
actually reaches close() instead of hanging in sess.join().
- pyOpenSSL is not thread-safe — reader-loop conn.* calls now hold
the same _send_lock the sender uses, dispatching decrypted packets
outside the lock so the auto-ACK send doesn't deadlock.
- Periodic CoAP Ping (RFC 7252 §4.4) keepalive to keep DTLS warm.
- Post-write Block2 fetchback REMOVED. It was the root cause of
every "setpoint/operationTime/modes revert ~3s after write"
symptom — Samsung's stack treats a read on a freshly-written
resource as a signal to invalidate that write. OBSERVE pushes
keep HA in sync without the verification GET.
HA-facing changes (oven):
- New entities: Lamp (light), Sound (switch), Fast preheat, Natural
steam, Setpoint (number), Cook time (number), Stop cycle (button).
- Cooking mode surfaced as a read-only sensor — the oven owns the
modes field once a cycle is active and rolls local writes back.
- Cook time writes operationTime + remainingTime on
/operational/state/vs/0 (discovered via OBSERVE capture of
SmartThings mid-cycle changes — UpperTimerSet on /mode/vs/0
options is vestigial and doesn't drive the running cycle).
- New cycle_active MQTT availability topic. Writes the oven only
honours mid-cycle (setpoint, cook time, fast preheat, natural
steam, stop) gate on it via avail_with_cycle / avail_with_remote_
and_cycle. Sound + Lamp remain always-available.
- Cycle Start deliberately NOT exposed. Every byte-level approxi-
mation of SmartThings's working start sequence is rejected at
the firmware level. Empty discovery payloads remove the previous
Start button and Cooking-mode select cleanly from HA.
Diagnostics:
- DEBUG_BRIDGE=1 env var enables verbose tracing (rx CON/NON/ACK/
RST per frame, full link-tree dump at seed, /oic/res directory,
REP changes on /operational/state, /oven, /power, mode options).
Quiet in production.