The refetch path logged only at debug, and the bridge configures logging
at INFO, so a successful re-read and a total failure produced identical
output: nothing. That makes the hardware validation for #39 impossible
to read.
One line per refetch, promoted to INFO when DEBUG_BRIDGE=1 and left at
debug otherwise, naming the href, the one-shot token, the block count,
and the reassembled size. The token is the part that matters: it is what
shows the re-read used a fresh 4-byte token rather than the observation's
1-byte one, which is the assumption the whole design rests on.
Gating on DEBUG_BRIDGE rather than raising the logger keeps the per-block
retransmit lines out of the way, and matches how the module already gates
its frame dump.
A notification carries only the first block of a large representation
(RFC 7959 §2.6). _dispatch_coap handed that block straight to
on_notification, so consumers decoded a truncated CBOR buffer. Reported
twice on /mode/vs/0: #37 and mbillow/localthings#361.
The recovery is a re-read from block 0 on a fresh 4-byte one-shot token,
not a §2.6 continuation. §3.4 rules out reusing the observation's token,
and this server drops a transfer that opens at NUM>0 under a token it
has not seen, so a continuation is the one shape that cannot work here.
A truncated notification is now withheld and queued to a worker thread
that re-reads the resource and delivers the reassembled representation.
When the re-read fails the notification is dropped at debug level and
the poll tiers carry freshness, which is what they already did.
The re-read has to run off the reader thread: _dispatch_coap runs there
and the transfer waits on an event only that same thread can set. The
worker is serialized and paces between transfers, so a notification
storm stays under the firmware request ceiling.
Also in the Block2 loop, now extracted and shared by both paths:
- compare the response's Block2 NUM against the one requested, so a
retransmitted block is no longer concatenated as if it were the next
- compare ETags across blocks (§2.4) and restart once when the
representation changes mid-transfer
- recompute the next block number from the accumulated byte offset when
the server negotiates the block size down
- re-check reader liveness while waiting on a block, so a mid-transfer
reader death fails fast instead of burning the whole timeout
- guard the token counters and _pending with a lock, now that the
session issues concurrent reads of its own
Closes#39
- add dtls_handshake.py (from #34) to the repo-layout tree
- point the issue #16 / #20 notes at SamsungServerProfile / ServerCertificateAuth instead of calling that path unsupported
- list the new certificate-profile, connect-deadline, and session-interruption test modules
The reader loop exited silently on any socket error, leaving conn/sock
set so the session still looked open. Every later get()/post()/ping()
then waited out its full request timeout on a session nobody was
reading, raising SessionTimeoutError on repeat, forever.
This started biting in v0.1.3 (d677c72), which moved to connected UDP
sockets: a connected socket surfaces ICMP errors on recv, so one
ECONNREFUSED from a rebooting appliance now killed the reader.
- Advisory ICMP errnos (ECONNREFUSED/EHOSTUNREACH/...) no longer kill
the reader; the next datagram usually works.
- Real reader exits log at WARNING; close()-driven exits stay quiet.
- A _reader_running Event lets get/post/ping/subscribe/refresh_observes
fail fast via _check_live() with SessionClosedError instead of waiting
out a timeout. Callers that never start a reader are unaffected.
Refs QuiteYellow/SmartThings-Local#37
The distribution checker enforces an exact-contents allowlist; add the
NOTICE file to the wheel dist-info/licenses expectation and the sdist
required set so the packaged trademark notice passes verification.
Add a Trademarks & disclaimer section to the README and a root NOTICE
file stating this is an independent, unofficial project not affiliated
with Samsung, and that Samsung/SmartThings marks are used nominatively.
Ship NOTICE inside the distributed artifacts by adding it to
license-files (wheel .dist-info/licenses/) and the sdist include list.
hatchling 1.31.0 ships only .gitignore in the sdist; older releases also
bundled .hgignore. check_sdist required an exact member set including
.hgignore, so the Validate workflow's package job failed on main once CI
resolved the newer hatchling.
Require the tracked source set plus the fixed metadata files, and accept
.gitignore/.hgignore as optional members either way.
- Note the Fedora/RHEL SHA-1 crypto-policy block in Part 2 and how
setup_cert.py auto-retries / the manual update-crypto-policies remedy.
- Add dtls_probe.py to the repo-layout tree (it was referenced in three
places but missing from the file listing).
- Drop the .venv/ prefix from the Part 1 probe command so it runs against
the pip-installed package before the Part 4 venv exists.
- Expand the tests parenthetical to name the probe, port-resolution, and
cert-signing suites.
* fix(setup_cert): surface openssl errors and work around SHA-1 crypto policy
The signing step forces -sha1 (the AC14K_M chain requires SHA-1-signed
leaves), which Fedora/RHEL's default crypto policy rejects on OpenSSL
3.x. run() also swallowed stderr, so the failure surfaced as an opaque
non-zero-exit traceback with no diagnostic.
- run() now raises CommandError carrying the command and openssl stderr
- mint_cert retries signing with a scoped OPENSSL_CONF enabling
rh-allow-sha1-signatures when the first attempt fails
- main() prints the update-crypto-policies fallback on failure
Fixes#15
* test(setup_cert): cover SHA-1 signing, error surfacing, and crypto-policy retry
Regression tests for the #15 fix:
- full mint_cert flow (SHA-1 leaf, UUID SAN, custom OIDs, chain assembly)
- CommandError surfaces openssl stderr on a genuine signing failure
- signing retries via the SHA-1 override when the plain attempt is blocked
- run() raises CommandError with detail
GitHub is deprecating the Node 20 action runtime; checkout@v4,
setup-python@v5, and upload/download-artifact@v4 all run on it and
were being auto-forced to Node 24 with a warning. Bump each to its
current major (checkout@v7, setup-python@v7, upload-artifact@v7,
download-artifact@v8), all of which run natively on Node 24.
Convert em-dash prose splices to varied punctuation (periods, colons,
semicolons, commas, parens), turn **Label.**-period bullets into
**Label:** colons, drop sentence-spanning bold in "Traps to avoid", and
cut a couple of hollow intensifiers.
No content, facts, tables, code, or links changed (50/50 line diff). Left
as-is: the `## Part N —` headings (heading-anchor stability), everything
inside code/log fences, table N/A cells, and numbered-list
`**Bold** — desc` carve-outs.
A stateless-by-default DTLS ClientHello probe that classifies a host:port
as DEAD/LIVE/COMPLETED/REJECTED in ~1 RTT off the server's first flight,
sitting in front of the full handshake.
Probe (smartthings_local/protocol/dtls_probe.py):
- Stateless liveness mode (default): stops at HelloVerifyRequest and never
sends the cookie'd second ClientHello, so by RFC 6347 §4.2.1 it leaves
no association on the device — safe to run before a real connect.
- Diagnostic mode (stateless=False): drives the handshake further to
capture cipher/cert-chain/CertificateRequest or a fatal Alert, for
OCF-PKI-wall characterization (#16). Kept out of hot reconnect paths.
- Retransmit + retries: services OpenSSL's DTLS retransmit timer so a
single dropped ClientHello no longer reads as a false DEAD.
MQTT bridge (mqtt_demo):
- Stateless pre-flight gate in session_once() rejects a silent/rebooting
device or wrong port in ~3s (retries=1) instead of eating the 12s
HANDSHAKE_TIMEOUT_S per reconnect.
- OCF-band port autodiscovery when OCF_PORT is unset: races the band in
parallel and returns on the first port to answer LIVE (~1 RTT, abandoning
the dead-port probes), cached across reconnects; the stateless gate
leaves no orphan, preserving the fixed-source-port §4.2.8 invariant.
Validated on real hardware (dryer 49155 / oven 49154): parallel discovery
resolves both ports in <1s, connect with no orphan cooldown, and a wrong
pinned port rejected in ~3s.
Tests: probe behaviour (retransmit recovery, stateless single-flight
guard, silent-port flight budget, diagnostic continuation) and bridge
port-resolution (pinned gate, parallel discovery early-exit, cache).
Root-cause fix for the stale-session stall on always-on appliances (#14).
When the client dies without close_notify (crash, SIGKILL), the device
keeps an orphaned DTLS association keyed to the old 5-tuple; a reconnect
from a fresh ephemeral port presents as a brand-new peer, so the orphan
lingers until the device's own timer reaps it (observed 5-15 min).
RFC 6347 §4.2.8 covers exactly this: a ClientHello arriving on an existing
association's 5-tuple means the peer rebooted, and the server must complete
the new handshake and discard the old association. Add an optional
local_port to DtlsCoapSession that binds the UDP source port, and have the
bridge bind base+appliance-index, so every reconnect re-handshakes over the
same 5-tuple and the orphan is evicted instead of waited out.
Bench-verified on live hardware (2026-07-26): RT-OCF accepts the
same-5-tuple rehandshake (oven, dryer: handshake completes over a
crash-orphaned association, reads work immediately). The oven does not
reproduce the fridge stall even with 11 crash-orphaned OBSERVE
registrations, so fridge-side confirmation of the eviction is still needed.
Co-authored-by: vmvarga <garrysuchiy@gmail.com>
- credit @indykoning + note localthings test path for the washer row
- soften DV90T mnid grouping (mnid=0AJT confirmed on DV5000T only)
- de-speculate the same-family note now that a washer is confirmed