refactor(registry): match board families by token table, not substring ladder

for_device_by_model() had grown to 21 sequential `if key is None` branches
and 102 comment lines against 59 lines of code -- 33 of the repo's 243
commits have touched this file. Most of that bulk came from one wrong
primitive: substring matching on a delimited string.

Samsung spells the same board family with either delimiter, so '_RAC_' and
'-RAC-' each needed their own rule, and 'ARTIK051_DONGLE_REF' (issues #77,
#83) matched no '_TOKEN_' spelling at all because REF lands at the end of
the pipe-prefix with no trailing underscore -- which is what
_model_num_segments() existed to work around. Which field a rule searched
(modelNum, or modelNum + description) was historical accident. Collisions
like WAC vs WA were resolved by one `if` physically preceding another,
invisible in the code and explained at length in prose.

Tokenize on any non-alphanumeric run, upper-case, and look the tokens up in
a flat table. Every delimiter spelling collapses to one entry, both fields
go through the same matcher in a documented order (modelNum, then
description, then the fuzzy consumer prefix), and specificity is a property
of the table rather than of line ordering.

Two behaviours are preserved deliberately:

- modelNum is matched before description, which is what keeps the legacy
  gas cooktop correct: it reports 'ARTIK051_GB_CT_001' (CT) alongside
  'ARTIK051_GLOBAL_COOKTOP' (COOKTOP, which otherwise means induction).
  It is the only known device whose two fields disagree.
- _consumer_model_key still splits on '_' only. Widening it to '-' would
  read the dishwasher's 'ADW-WW-RTL-24-AILITE' board segment as a bare 'WW'
  washer.

Verified identical: all 49 device fixtures resolve to the same registry
before and after, and every existing for_device_by_model test case passes
unchanged. The table also picks up two families that previously depended on
oneUiVersion alone (TP1X_DA-AC-AIR air purifiers, ADW dishwashers), so they
now survive firmware that omits it.

TestBoardTokenAmbiguity guards the one property the flat lookup needs --
that no real model string contains two tokens naming different device types
-- across the whole fixture corpus, so a newly added dump exercises it
automatically.

The skill gains a section on routing: what each detection stage is for, the
rules for adding a token (name the specific type, never the board family;
never add a delimiter spelling; two-letter tokens are a last resort), when
to reach for the consumer prefix or a resource signature instead, and the
measured stake -- an unrouted device loses roughly half its entities.
This commit is contained in:
Marc Billow
2026-07-29 01:57:10 +00:00
parent 288ba02cc5
commit a863df1e59
4 changed files with 342 additions and 185 deletions
+103 -10
View File
@@ -5,6 +5,8 @@ description: >-
/device/0 diagnostics dump. Use when a device-support issue lands, a device
raises the "incomplete capability coverage" repair, a diagnostics JSON needs
triaging, or you're mapping OCF resources to HA entities. Covers reading dumps,
routing an unrecognized board family to a registry (oneUiVersion, modelNum
board tokens, resource signatures),
OCF-standard vs vendor hrefs, the diagnostic/config/normal entity taxonomy,
preferring dynamic (device-reported) select options over hardcoded lists,
ensuring every href is bound or ignored, and locking it in with a fixture +
@@ -47,9 +49,15 @@ discovery = importlib.import_module('custom_components.localthings.registry.disc
adapter = importlib.import_module('custom_components.localthings.registry.adapter')
resources = json.load(open('dump.json'))['data']['resources']
info = resources['/information/vs/0']
reg = by_type.for_device_by_model(info['x.com.samsung.da.modelNum'], info['x.com.samsung.da.description'])
# or: by_type.for_device(one_ui_version) when /otninformation has swVersionInfo.oneUiVersion
info = resources.get('/information/vs/0', {})
one_ui = resources.get('/otninformation/vs/0', {}).get('swVersionInfo', {}).get('oneUiVersion', '')
# Same three-stage order the coordinator uses -- see §3.
reg = (
(by_type.for_device(one_ui) if one_ui else None)
or by_type.for_device_by_model(info.get('x.com.samsung.da.modelNum', ''),
info.get('x.com.samsung.da.description', ''))
or by_type.for_device_by_resources(resources)
)
unbound = []
bound = discovery.discover(resources, reg.capabilities, reg.pattern_capabilities, log=unbound.append)
state = adapter.flatten(bound, resources) # {entity_key: value}
@@ -61,7 +69,86 @@ print('state_keys:', sorted(state))
`exists_fn` and produces the final entity values. Use the same routine to
regenerate a golden.
## 3. OCF-standard vs vendor hrefs (`/x/0` vs `/x/vs/0`)
## 3. Route the device to a registry — add a row, never a branch
If `for_device*` returns `None`, the device falls back to common capabilities
and loses roughly **half** its entities (measured across the fixture corpus:
843 of 1510 bound entities survive). So routing is the first thing to fix, and
`registry/by_type/__init__.py` is deliberately kept boring:
1. **`for_device(one_ui_version)`** — `/otninformation/vs/0`'s
`swVersionInfo.oneUiVersion`, e.g. `'7.0 Dishwasher'`. The device naming
its own type, so it's tried first — but only a minority of hardware
reports it, so never assume it exists.
2. **`for_device_by_model(model_num, description)`** — the workhorse. Both
fields come from `/information/vs/0`. Board-family tokens are matched
against `modelNum` first, then `description`, then the fuzzy two-letter
consumer-model prefix.
3. **`for_device_by_resources(resources)`** — last resort for boards that
report no `/information/vs/0` at all. Needs a *distinctive* signature.
### Adding a board family
Almost always a one-line addition to `_BOARD_TOKEN_TO_KEY`:
```python
'VSKR': 'vacuum_station', # issue #131 -- stick-vacuum clean station
```
Matching is on **whole tokens** of the model string, split on any run of
non-alphanumerics and upper-cased. That is what keeps this a table, and it
carries rules:
- **Never add a delimiter spelling.** `'_RAC_'` and `'-RAC-'` are the same
entry, `RAC`. If you find yourself adding a second row for punctuation, the
tokenizer already handled it.
- **Name the specific type, never the board family.** `DA-AC-` prefixes
RAC/WAC/DHM/AIR alike — a bare `'AC'` row would swallow the dehumidifier
and the air purifier. Same for `DA`, `KS`, `WM`, `TP1X`, `ARTIK051`.
`TestBoardTokenTable` asserts these stay out.
- **Never add a token that can co-occur with another.** `_board_family_key`
returns the first hit, which is only safe while no real model string
contains two tokens naming different types.
`TestBoardTokenAmbiguity` checks that invariant against every fixture, so a
new dump exercises it automatically — if it fails, the answer is a narrower
token, not a reordering.
- **Two-letter tokens are a last resort.** `'CT'` (legacy gas cooktop) is the
only one, and it is loose enough to collide by accident.
Reach past the table only when the evidence isn't a board token:
- **Consumer-model prefix** (`_CONSUMER_PREFIX_TO_KEY`) — for washers, dryers
and dishwashers, whose `modelNum` is the shared `DA_WM_` laundry board and
whose real type is only in `description`'s trailing model code
(`..._WA8000T`). Deliberately split on `_` only: widening it to `-` would
read the dishwasher's `ADW-WW-RTL-24-AILITE` board segment as a `WW`
washer. Consulted last because a two-letter prefix is the weakest evidence
here — `WAC` (window AC) starts with `WA` (top-load washer).
- **Resource signature** (`for_device_by_resources`) — only when
`/information/vs/0` is absent entirely. Require **two** independent shapes
(e.g. `/oven/vs/0` present *and* a `MicroWave*` entry in `supportedModes`),
never one, or an unrelated family's `/mode/vs/0` will match.
### If the model string identifies nothing
Check the diagnostics `identity` block before inventing a rule: it carries
`/oic/p` and `/oic/d`, which sit outside the `/device/0` dump.
`identity.device_types` is `/oic/d`'s `rt` — OCF's own device-type
declaration (`oic.d.airconditioner`). Nothing routes on it yet because no
captured dump has ever included it; if real hardware turns out to populate it,
it beats parsing board part numbers and this whole section shrinks. Note in
the issue when a dump has it.
### Sharing a registry vs adding one
Route a new family to an **existing** registry when its resource surface
matches (most AC board families do — verify by checking the dump binds with
zero unbound hrefs). Add a **new** registry only when the resources genuinely
differ: `vacuum_station` earned one because it shares no hrefs with anything
modelled; `microwave` split from `oven` over a distinct mode vocabulary,
setpoint bounds, and a `powerLevel` field.
## 4. OCF-standard vs vendor hrefs (`/x/0` vs `/x/vs/0`)
Samsung appliances run RT-OCF and often expose the **same state twice**:
- `/x/vs/0` — **vendor** resource, `x.com.samsung.da.*` fields.
@@ -83,7 +170,7 @@ resource from the populated dump:
Course/cycle is **not** an OCF question — there's no standard course resource, so
`/course/vs/0` (and the `/st/*course/vs/0` re-encoding) are both vendor.
## 4. Entity taxonomy — the judgement call
## 5. Entity taxonomy — the judgement call
For each field worth exposing, decide the entity kind and category
(`entity_category` on the descriptor):
@@ -104,7 +191,7 @@ sub-polled between summary polls. Pick descriptor types from `entities.py`
as a gap for a human, or ignore it with a documented reason — never invent an
entity on a hunch (`ignored.py`'s rule).
## 5. Select options: read them from the device, don't hardcode
## 6. Select options: read them from the device, don't hardcode
A `SelectDesc`'s `options` should come from the device's own advertised list
whenever the resource carries one, not from a Python tuple typed in from a
@@ -134,7 +221,7 @@ permanent design choice: migrate it to `options_field`/a callable the moment
a dump with a real supported-values list surfaces, instead of just adding
the new values to the static tuple.
## 6. Names and enum labels live in translations, never in Python
## 7. Names and enum labels live in translations, never in Python
Descriptors have **no `name` field**. Every entity is named from the shipped
catalog, keyed by `translation_key` — which defaults to the descriptor's own
@@ -173,7 +260,7 @@ no `[%key:...%]` resolution (that's Core build tooling). Every other language
must mirror `en.json` key for key — also enforced by
`tests/test_translations.py`.
## 7. Coverage discipline: bound or ignored
## 8. Coverage discipline: bound or ignored
Every href in the dump must resolve, or the repair fires. If a resource isn't
worth an entity, add it to `capabilities/ignored.py` (a no-entity `Capability`)
@@ -187,7 +274,7 @@ friendlier href**.
ignored because washers bind it. When only one family should ignore an href
that another binds, scope the ignore to that family's registry.
## 8. Reuse before writing new code
## 9. Reuse before writing new code
Check `common.py` (generic OCF: power, energy, alarms, water) and `laundry.py`
(shared washer/dryer/dishwasher: buzzer, job status, `cycle_select` + course
@@ -196,7 +283,7 @@ registry uses `fridge.FIRMWARE_UPDATE`; all three laundry families share
`laundry.cycle_select`. If two families hand-roll the same helper, hoist it to a
shared module rather than copying.
## 9. Lock it in
## 10. Lock it in
1. Add a **scrubbed** fixture `tests/fixtures/<type>_device.json`
(`{"device0": [ {devcol rep}, {href, rep}, ... ]}`) — replace serials, MACs,
@@ -209,6 +296,12 @@ shared module rather than copying.
4. Run `pytest tests/ -q` — and re-run the golden tests for **other** device
types after any change to `common.py`/`laundry.py`, since they share those.
The new fixture is picked up automatically by the corpus-wide checks (the
`all_device_fixtures` conftest fixture), including
`TestBoardTokenAmbiguity` — so a model string that collides with an existing
board token fails the build rather than silently mistyping someone's
appliance.
## Key files
- `registry/discovery.py` — `discover()`, unbound reporting, pattern caps.
- `registry/capability.py`, `registry/entities.py` — the `Capability` and
@@ -1,4 +1,5 @@
"""Per-device-type registries."""
import re
from typing import Optional
from ._base import DeviceRegistry
@@ -10,7 +11,7 @@ from . import (
__all__ = [
'DeviceRegistry', '_type_key', 'for_device', 'for_device_by_model',
'for_device_by_resources',
'for_device_by_resources', '_board_tokens',
]
@@ -94,21 +95,89 @@ _CONSUMER_PREFIX_TO_KEY: dict[str, str] = {
'DW': 'dishwasher',
}
# Board-family token -> registry key, matched against whole tokens of
# `modelNum`/`description` (see `_board_tokens`).
#
# Tokenizing instead of substring-matching is what keeps this a table rather
# than a ladder of hand-written rules. Samsung spells the same board family
# with either delimiter -- 'TP1X_DA-AC-RAC-01001' and 'TP2X_RAC_20K' are the
# same RAC family -- so a substring rule has to be written once per spelling
# ('_RAC_' *and* '-RAC-'), and a token that lands at the end of the
# pipe-prefix with no trailing delimiter ('ARTIK051_DONGLE_REF', issues #77
# and #83) matches no '_TOKEN_' spelling at all. Whole-token matching sees
# every one of those as a single entry.
#
# Entries must name the *specific* device type, never the board family that
# contains it: 'DA-AC-' prefixes RAC/WAC/DHM/AIR alike, so a bare 'AC' entry
# would swallow the dehumidifier and the air purifier. Where two families
# genuinely share a resource surface they share a registry (all the
# air-conditioner spellings below), which is a statement about the hardware,
# not a shortcut.
_BOARD_TOKEN_TO_KEY: dict[str, str] = {
'REF': 'refrigerator',
# Air conditioners. Every one of these is a distinct board family with
# the same resource surface: room (issues #37, #91), package, Korean
# (#136), window (#87), 2-in-1 floor+wall (#150, #153), system/commercial
# (#52), and ARA-WW wall-mount (#115, #116, #117, #120).
'RAC': 'airconditioner',
'PRAC': 'airconditioner',
'KRAC': 'airconditioner',
'WAC': 'airconditioner',
'FAC': 'airconditioner',
'CAWW': 'airconditioner',
'ARA': 'airconditioner',
'DHM': 'dehumidifier', # issue #88 -- target humidity, no climate
'TVTL': 'air_purifier', # issue #56 (ARTIK051)
'VTWW': 'air_purifier', # issue #151 (BESPOKE Cube Air)
'AIR': 'air_purifier', # issue #130 (TP1X_DA-AC-AIR)
'WATERPURIFIER': 'water_purifier', # issue #90
'ADW': 'dishwasher',
'AHD': 'range_hood',
'RANGE': 'range', # issue #44 -- cooktop+oven combo
'OVEN': 'oven', # issue #55 -- wall oven, no burners
'MICROWAVE': 'microwave', # issues #66, #121
'COOKTOP': 'induction_cooktop', # issue #86 -- standalone, no oven
# Legacy ARTIK051 gas cooktops ('ARTIK051_GB_CT_001'), whose burner state
# lives in /mode/vs/0's options array. Deliberately a bare two-letter
# token, and so the loosest entry in this table -- it is only ever
# reached by a device that matched nothing more specific, and its
# `description` ('ARTIK051_GLOBAL_COOKTOP') would otherwise read as an
# induction cooktop via the COOKTOP entry above. See `for_device_by_model`
# for the field ordering that makes that resolve correctly.
'CT': 'cooktop',
'VSKR': 'vacuum_station', # issue #131 -- stick-vacuum clean station
'DF': 'air_dresser', # issue #162
}
def _model_num_segments(model_num: str) -> list[str]:
"""Underscore-delimited segments of modelNum's pipe-prefix.
_TOKEN_SPLIT_RE = re.compile(r'[^A-Z0-9]+')
Most boards wrap a token in underscores on both sides ('..._REF_...'),
which a plain substring check catches fine. But the ARTIK051_DONGLE_REF
family (issues #77, #83) reports modelNum as
'<board>_DONGLE_REF|<rest...>' -- REF is the *last* segment before the
pipe, with no trailing underscore, so '_REF_' never matches and the
device silently fell back to 'unknown'. Splitting on '_' and checking
segment membership catches both shapes without caring which side (if
either) has a delimiter.
def _board_tokens(value: str, cut_at: str) -> list[str]:
"""Whole, upper-cased tokens of `value` up to the first `cut_at`.
`cut_at` drops the trailing junk each field carries -- everything after
modelNum's first '|' (a board revision and a capability bitmap, which can
contain anything) and after description's first '/' (a '/DC92-...' board
part number).
"""
prefix = (model_num or '').split('|', 1)[0]
return prefix.split('_')
head = (value or '').split(cut_at, 1)[0].upper()
return [t for t in _TOKEN_SPLIT_RE.split(head) if t]
def _board_family_key(value: str, cut_at: str) -> Optional[str]:
"""First `_BOARD_TOKEN_TO_KEY` hit among `value`'s tokens, or None.
No known modelNum or description yields two *conflicting* board keys, so
which token is found first doesn't matter within one field -- the table is
a flat lookup, not a priority list. Adding an entry that could co-occur
with another (a family token, or one short enough to collide by accident)
would break that property; see this table's comment.
"""
for token in _board_tokens(value, cut_at):
key = _BOARD_TOKEN_TO_KEY.get(token)
if key is not None:
return key
return None
def _consumer_model_key(description: str) -> Optional[str]:
@@ -123,12 +192,18 @@ def _consumer_model_key(description: str) -> Optional[str]:
segments from the end and take the first one that resolves, rather
than assuming the last segment is always it.
Splits on '_' only, unlike `_board_tokens` above: these are two-letter
prefixes matched against the *start* of a segment, so widening the split
to '-' as well would start reading board-family segments as consumer
models -- the dishwasher's 'ADW-WW-RTL-24-AILITE' would offer up a bare
'WW' segment and route to washer.
Only a 2-letter *prefix* match -- e.g. 'WAC' (the Window Air Conditioner
board-family token, issue #87) also starts with 'WA' (the top-load-washer
prefix, issue #106) at this granularity. for_device_by_model() calls the
board-family modelNum checks first and this function only as a fallback,
so that ambiguity resolves correctly without this function needing to
know about unrelated device families.
prefix, issue #106) at this granularity. for_device_by_model() consults
the board-family table first and this function only as a fallback, so
that ambiguity resolves correctly without this function needing to know
about unrelated device families.
"""
segments = (description or '').split('/', 1)[0].split('_')
for segment in reversed(segments):
@@ -143,161 +218,34 @@ def for_device_by_model(model_num: str, description: str) -> Optional[DeviceRegi
oneUiVersion (confirmed for washers -- their /otninformation/vs/0 has
no swVersionInfo key at all).
Three passes, narrowest evidence first:
1. Board-family tokens in `modelNum`. The most reliable signal -- it names
the board, which determines the resource surface.
2. The same tokens in `description`. Some units carry the board token only
there (a scrubbed or placeholder modelNum, e.g. description
'TP1X_REF_21K'). This runs second so that a device whose two fields
disagree is typed by its modelNum: the legacy gas cooktop reports
'ARTIK051_GB_CT_001' (CT -> gas cooktop) alongside
'ARTIK051_GLOBAL_COOKTOP' (COOKTOP -> induction cooktop), and the
board is right.
3. The consumer-model prefix in `description` (washer/dryer/dishwasher).
Last, because a bare two-letter prefix is the fuzziest evidence here
and would otherwise shadow the specific board tokens above.
Args:
model_num: x.com.samsung.da.modelNum from /information/vs/0.
description: x.com.samsung.da.description from /information/vs/0.
Returns:
DeviceRegistry if the consumer-model code or modelNum resolves to a
DeviceRegistry if the modelNum or consumer-model code resolves to a
known type, None otherwise.
"""
# Board-family modelNum tokens are checked first, and the fuzzier
# 2-letter consumer-model prefix from `description` only as a fallback
# (see the bottom of this function) -- some board tokens are themselves
# only 2-3 letters long ('WAC', issue #87) and would otherwise collide
# with an unrelated consumer prefix at that granularity ('WA', issue #106).
key = None
if key is None and 'REF' in _model_num_segments(model_num):
key = 'refrigerator'
# Room air conditioners (e.g. ARTIK051_PRAC_20K) report no oneUiVersion and
# a modelNum carrying the '_PRAC_' (Package Room Air Conditioner) token.
if key is None and '_PRAC_' in (model_num or ''):
key = 'airconditioner'
# Other RAC boards carry a bare 'RAC' (Room Air Conditioner) token in the
# modelNum, in one of two spellings: the underscore form '_RAC_' (e.g.
# TP2X_RAC_20K, issue #37) or the hyphenated form '-RAC-' (e.g.
# TP1X_DA-AC-RAC-01001, a cool-only global variant, issue #91). Both are
# distinct from '_PRAC_' above ('P' sits before 'RAC' with no delimiter)
# and from range ('-RANGE-') / oven ('-OVEN-') tokens. Most TP1X boards
# self-report oneUiVersion and resolve via for_device() upstream of this
# fallback; the hyphenated match is what rescues variants whose
# /otninformation/vs/0 ships no swVersionInfo block at all, so
# oneUiVersion is empty.
if key is None and ('_RAC_' in (model_num or '')
or '-RAC-' in (model_num or '').upper()):
key = 'airconditioner'
# System air conditioners (multi-indoor-unit commercial installs, e.g.
# A-CAWW-TP2-20-COMMON, issue #52) report no oneUiVersion either and
# carry the '-CAWW-' board-family token instead of '_RAC_'/'_PRAC_'.
# Same TP1X/TP2X-class resource surface as the room-AC models above
# (confirmed by the issue #52 dump binding cleanly against the existing
# airconditioner registry once routed here), plus one new SAC-specific
# resource (see airconditioner.py's _AC_IGNORED).
if key is None and '-CAWW-' in (model_num or '').upper():
key = 'airconditioner'
# Dehumidifiers (e.g. TP1X_DA_AC_DHM_01001_0000, issue #88) share the
# DA_AC_ board family with the room-AC models above but carry the
# '_DHM_' (DeHuMidifier) token instead of '_RAC_'/'_PRAC_'. Distinct
# registry: target humidity, not temperature, is the primary control,
# and there's no climate composite.
if key is None and '_DHM_' in (model_num or ''):
key = 'dehumidifier'
# Window air conditioners (e.g. TP1X_DA_AC_WAC_01001_0000, issue #87)
# report no oneUiVersion and carry the '_WAC_' (Window Air Conditioner)
# token instead of '_RAC_'/'_PRAC_'. Same TP1X-class resource surface
# as the room-AC models above (mode/convenient/wind/temperature/power/
# filter/humidity all confirmed against the issue #87 dump binding
# cleanly against the existing airconditioner registry once routed
# here) -- no WAC-specific resources needed.
if key is None and '_WAC_' in (model_num or ''):
key = 'airconditioner'
# Wind-Free 2-in-1 systems (floor-standing + wall-mounted indoor units
# sharing one outdoor unit and one local IP, e.g. TP2X_FAC_BORA_21K,
# issues #150/#153) report no oneUiVersion and carry the '_FAC_' token
# instead of '_RAC_'/'_PRAC_'. Same TP1X/TP2X-class resource surface as
# the room-AC models above.
if key is None and '_FAC_' in (model_num or ''):
key = 'airconditioner'
# ARA-WW-class wall-mount RACs (e.g. ARA-WW-TP1-22-COMMON, issues #115/
# #116/#117/#120) report no oneUiVersion and no '_RAC_'/'-RAC-' token at
# all -- the board family is spelled 'ARA-WW-' instead. Same resource
# surface as the other TP1X-class room ACs above (mode/convenient/wind/
# temperature/power/filter/humidity all confirmed against these dumps
# binding cleanly against the existing airconditioner registry once
# routed here) -- no ARA-WW-specific resources needed, so this reuses
# the same registry rather than adding a new device type.
if key is None and 'ARA-WW-' in (model_num or '').upper():
key = 'airconditioner'
# Air purifiers (e.g. ARTIK051_TVTL_18K, issue #56) report no
# oneUiVersion either, and carry the '_TVTL_' board-family token.
# Room air conditioners on the ARTIK051 board (e.g. ARTIK051_KRAC_18K,
# issue #136) report no oneUiVersion and carry a '_KRAC_' token. The '_RAC_'
# check above cannot see it -- the 'K' sits between the underscore and 'RAC' --
# and the consumer-prefix fallback only covers washers/dryers/dishwashers, so
# these units fell back to 'unknown' and exposed nothing but power. Same
# ARTIK051 board family as the '_TVTL_' air purifier below.
if key is None and '_KRAC_' in (model_num or ''):
key = 'airconditioner'
if key is None and '_TVTL_' in (model_num or ''):
key = 'air_purifier'
# BESPOKE Cube Air (e.g. A-VTWW-TP2-21-COMMON, issue #151) reports no
# oneUiVersion and carries the hyphenated '-VTWW-' board-family token
# (distinct from the underscore-delimited '_TVTL_' ARTIK051 family
# above). Its fan lives on /wind/strength/vs/0 rather than /mode/vs/0 --
# see capabilities/air_purifier.py's WIND_STRENGTH_FAN.
if key is None and '-VTWW-' in (model_num or '').upper():
key = 'air_purifier'
model_identity = f'{model_num} {description}'.upper()
if key is None and ('_COOKTOP' in model_identity or '_GB_CT_' in model_identity):
key = 'cooktop'
# Water purifiers (e.g. TP2X_WATERPURIFIER_20K, issue #90) report no
# oneUiVersion and carry the 'WATERPURIFIER' board-family token.
if key is None and 'WATERPURIFIER' in model_identity:
key = 'water_purifier'
if key is None and model_identity.startswith('AHD-'):
key = 'range_hood'
# Range/cooktop-oven combos (e.g. TP1X_DA-KS-RANGE-0102X, issue #44) --
# like the RAC/PRAC air conditioners above, these report no oneUiVersion
# and don't match the washer/dryer/dishwasher consumer-prefix map either.
if key is None and '-RANGE-' in (model_num or '').upper():
key = 'range'
# Wall ovens (e.g. TP1X_DA-KS-OVEN-0107X, issue #55) -- same board-family
# naming as the range combo above, minus the burners; also reports no
# oneUiVersion and doesn't match the washer/dryer/dishwasher prefix map.
if key is None and '-OVEN-' in (model_num or '').upper():
key = 'oven'
# Microwaves, both combi (e.g. TP1X_DA-KS-MICROWAVE-01041, issue #121)
# and plain (e.g. TP2X_DA-KS-MICROWAVE-01011, issue #66) -- same board
# family as the wall oven above (an '/oven/vs/0' cavity resource, same
# /operational/state/vs/0 + /doors/vs/0 shape), but a distinct mode
# vocabulary (Convection/AirFryer/Grill/MicroWave*) and setpoint bounds
# from the oven registry, plus a powerLevel field ovens don't report --
# its own device type rather than folded into 'oven' (issue #121 shipped
# it onto the oven registry initially; split out per user feedback).
# Reports no oneUiVersion and doesn't match the washer/dryer/dishwasher
# prefix map either.
if key is None and '-MICROWAVE-' in (model_num or '').upper():
key = 'microwave'
# Standalone induction cooktops (e.g. TP1X_DA-KS-COOKTOP-01011, issue
# #86) -- same board family and '/cooktop/status/vs/0' resource shape
# as the range combo above, but no oven attached at all. Distinct from
# the '_COOKTOP'/'_GB_CT_' underscore-delimited check above: that one
# matches a different, older gas-cooktop family (cooktop.py's NA9300K
# class, burner state embedded in /mode/vs/0's options array) whose
# modelNum token is underscore-delimited, not hyphenated like this
# board family's.
if key is None and '-COOKTOP-' in (model_num or '').upper():
key = 'induction_cooktop'
# Stick-vacuum clean/auto-empty station (e.g. A-VSKR-TP1-22-VS9500AL,
# issue #131) -- reports no oneUiVersion and carries the '-VSKR-'
# board-family token. See capabilities/vacuum_station.py for why this
# is its own device type: the resource set (dustbag/dustbin/UV-sanitize
# station state) shares no hrefs with anything else already modeled.
if key is None and '-VSKR-' in (model_num or '').upper():
key = 'vacuum_station'
# AirDresser (e.g. DA_DF_A51_20_COMMON, issue #162) -- reports no
# oneUiVersion and carries the '_DF_' (Dresser Function) board-family
# token. Every resource it exposes (course select, wrinkle-prevent
# setting, diagnosis) is already handled by the shared laundry
# machinery; it needs its own device type only because none of the
# washer/dryer/dishwasher consumer-model prefixes below match it.
if key is None and '_DF_' in (model_num or ''):
key = 'air_dresser'
# Consumer-model prefix from `description` (washer/dryer/dishwasher) --
# last, since it's the fuzziest match (a bare 2-letter prefix) and would
# otherwise shadow the more specific board-family tokens above.
if key is None:
key = _consumer_model_key(description)
key = (
_board_family_key(model_num, '|')
or _board_family_key(description, '/')
or _consumer_model_key(description)
)
return _REGISTRY_BY_KEY.get(key) if key else None
+15
View File
@@ -41,3 +41,18 @@ def fridge_resources() -> dict[str, dict]:
@pytest.fixture
def washer_resources() -> dict[str, dict]:
return _load_device('washer')
@pytest.fixture
def all_device_fixtures() -> dict[str, dict[str, dict]]:
"""Every scrubbed device dump, keyed by fixture name.
For invariants that must hold across the whole corpus rather than for one
device -- so a newly added dump exercises them automatically.
"""
return {
path.name[:-len('_device.json')]: _resources_from_dump(
json.loads(path.read_text())
)
for path in sorted(FIXTURES.glob('*_device.json'))
}
+111 -10
View File
@@ -114,23 +114,60 @@ class TestWasherRegistry:
assert href in registry.capabilities, f"{href} missing from washer registry"
class TestModelNumSegments:
def test_splits_pipe_prefix_on_underscore(self):
from custom_components.localthings.registry.by_type import _model_num_segments
assert _model_num_segments('ARTIK051_DONGLE_REF|00127641|000800200014') == [
class TestBoardTokens:
def test_splits_pipe_prefix_into_whole_tokens(self):
from custom_components.localthings.registry.by_type import _board_tokens
assert _board_tokens('ARTIK051_DONGLE_REF|00127641|000800200014', '|') == [
'ARTIK051', 'DONGLE', 'REF',
]
def test_ignores_everything_after_first_pipe(self):
from custom_components.localthings.registry.by_type import _model_num_segments
assert _model_num_segments('TP2X_RAC_20K|abc|REF_should_not_appear') == [
def test_ignores_everything_after_the_cut(self):
from custom_components.localthings.registry.by_type import _board_tokens
assert _board_tokens('TP2X_RAC_20K|abc|REF_should_not_appear', '|') == [
'TP2X', 'RAC', '20K',
]
def test_both_delimiters_produce_the_same_tokens(self):
"""The whole point of tokenizing: Samsung spells one board family
with either delimiter, and both must reduce to the same tokens."""
from custom_components.localthings.registry.by_type import _board_tokens
assert (_board_tokens('TP1X_DA-AC-RAC-01001_0000', '|')
== _board_tokens('TP1X_DA_AC_RAC_01001_0000', '|')
== ['TP1X', 'DA', 'AC', 'RAC', '01001', '0000'])
def test_upper_cases_and_drops_empty_runs(self):
from custom_components.localthings.registry.by_type import _board_tokens
assert _board_tokens('a--b__c', '|') == ['A', 'B', 'C']
def test_empty_for_none_or_empty_input(self):
from custom_components.localthings.registry.by_type import _model_num_segments
assert _model_num_segments('') == ['']
assert _model_num_segments(None) == ['']
from custom_components.localthings.registry.by_type import _board_tokens
assert _board_tokens('', '|') == []
assert _board_tokens(None, '|') == []
class TestBoardTokenTable:
def test_no_board_family_token_shadows_a_specific_type(self):
"""'DA-AC-' prefixes RAC/WAC/DHM/AIR alike -- a bare 'AC' entry would
type the dehumidifier and the air purifier as air conditioners."""
from custom_components.localthings.registry.by_type import _BOARD_TOKEN_TO_KEY
for family_token in ('AC', 'DA', 'KS', 'WM', 'TP1X', 'TP2X', 'ARTIK051'):
assert family_token not in _BOARD_TOKEN_TO_KEY
def test_every_token_resolves_to_a_real_registry(self):
from custom_components.localthings.registry.by_type import (
_BOARD_TOKEN_TO_KEY, _CONSUMER_PREFIX_TO_KEY, _REGISTRY_BY_KEY,
)
for token, key in _BOARD_TOKEN_TO_KEY.items():
assert key in _REGISTRY_BY_KEY, f"{token!r} -> unknown registry {key!r}"
for prefix, key in _CONSUMER_PREFIX_TO_KEY.items():
assert key in _REGISTRY_BY_KEY, f"{prefix!r} -> unknown registry {key!r}"
def test_tokens_are_upper_case(self):
"""`_board_tokens` upper-cases before lookup, so a lower-case entry
would be dead."""
from custom_components.localthings.registry.by_type import _BOARD_TOKEN_TO_KEY
for token in _BOARD_TOKEN_TO_KEY:
assert token == token.upper()
class TestConsumerModelKey:
@@ -452,6 +489,70 @@ class TestForDeviceByModel:
from custom_components.localthings.registry.by_type import for_device_by_model
assert for_device_by_model('', '') is None
@pytest.mark.parametrize('model_num', [
'TP1X_DA-AC-RAC-01001_0000', # hyphenated (issue #91)
'TP1X_DA_AC_RAC_01001_0000', # underscored
'TP2X_RAC_20K', # bare, no board-family prefix (issue #37)
'TP2X-RAC-20K',
])
def test_delimiter_spelling_does_not_change_the_answer(self, model_num):
"""One board family, four spellings, one table entry."""
from custom_components.localthings.registry.by_type import for_device_by_model
reg = for_device_by_model(model_num, '')
assert reg is not None
assert reg.name == 'airconditioner'
def test_model_num_wins_when_the_two_fields_disagree(self):
"""The legacy gas cooktop is the one known device whose fields
conflict: modelNum says CT (gas), description says COOKTOP (which
otherwise means induction). The board is right, so modelNum is
consulted first."""
from custom_components.localthings.registry.by_type import for_device_by_model
reg = for_device_by_model('ARTIK051_GB_CT_001', 'ARTIK051_GLOBAL_COOKTOP')
assert reg is not None
assert reg.name == 'gas_cooktop'
def test_board_token_in_description_used_when_model_num_has_none(self):
"""Some units report a placeholder modelNum and carry the board token
only in `description`."""
from custom_components.localthings.registry.by_type import for_device_by_model
reg = for_device_by_model('TEST-MODEL', 'TP1X_REF_21K')
assert reg is not None
assert reg.name == 'refrigerator'
def test_board_token_beats_consumer_prefix(self):
"""'WAC' (window AC, issue #87) starts with 'WA' (top-load washer,
issue #106). The board table runs first, so the AC wins."""
from custom_components.localthings.registry.by_type import for_device_by_model
reg = for_device_by_model('TP1X_DA_AC_WAC_01001_0000', 'TP1X_DA_AC_WAC_01001_0000')
assert reg is not None
assert reg.name == 'airconditioner'
class TestBoardTokenAmbiguity:
"""`_board_family_key` returns the first matching token, which is only
safe while no real model string contains two tokens naming different
device types. Guard that against every dump we have."""
def test_no_fixture_model_string_yields_two_conflicting_keys(self, all_device_fixtures):
from custom_components.localthings.registry.by_type import (
_BOARD_TOKEN_TO_KEY, _board_tokens,
)
for name, resources in all_device_fixtures.items():
info = resources.get('/information/vs/0', {})
for field, cut in (
(info.get('x.com.samsung.da.modelNum', ''), '|'),
(info.get('x.com.samsung.da.description', ''), '/'),
):
keys = {
_BOARD_TOKEN_TO_KEY[t]
for t in _board_tokens(field, cut)
if t in _BOARD_TOKEN_TO_KEY
}
assert len(keys) <= 1, (
f"{name}: {field!r} matches conflicting board tokens {keys}"
)
class TestForDeviceByResources:
def test_na9300k_without_one_ui_or_information_is_cooktop(self):