Commit Graph
35 Commits
Author SHA1 Message Date
Timothy Jaeryang Baek 45e49d33e5 refac 2026-04-13 21:52:19 -05:00
Timothy Jaeryang Baek 25898116ea chore: format 2026-04-12 18:12:59 -05:00
Timothy Jaeryang Baek 15b89b9218 refac 2026-04-12 16:56:00 -05:00
Algorithm5838 b6db719758 perf: build mention regex once in factory closure (#23551) 2026-04-11 15:02:17 -06:00
Timothy Jaeryang Baek eb5c95ef8e refac 2026-04-01 05:30:07 -05:00
Timothy Jaeryang Baek 53b8a1f71b enh: colon fence md 2026-03-21 17:23:38 -05:00
Classic298 8cd2157564 Perf: precompile katex unicode regex (#22196)
* perf: pre-compile KaTeX Unicode regex at module load time

The katexStart() function was creating a new RegExp with Unicode
property escapes (\p{Script=Han}, \p{Script=Hiragana}, etc.) on
every invocation. Unicode property escapes are extremely expensive
to compile as the regex engine must build character class tables
covering tens of thousands of code points.

Since marked calls the start() function at every character position
while scanning source text, this meant hundreds of regex compilations
per marked.lexer() call, and lexer runs ~60 times/sec during streaming.
Profiling showed KaTeX regex consuming 87% (320ms/365ms) of total
markdown rendering time.

Changes:
- Pre-compile SURROUNDING_CHARS_REGEX once at module load time
- Use .test() instead of .match() to avoid array allocations
- Fix delimiter search to find earliest match, not last match

* perf: replace katexStart with single-pass character scan

The katexStart() function was the dominant cost in marked's lexer,
consuming 55-58% of total markdown rendering time per profiling.

It was called at every character position by marked and each call:
- Looped through 3-5 delimiters, each doing indexOf() on the full
  remaining source (3-5 x O(n) string scans per call)
- Ran the complex ruleReg regex with Unicode lookaheads for validation
- On failed validation, created substrings and looped again

Replace with a single linear character scan using charCodeAt that:
- Checks only for $ (charCode 36) or backslash (charCode 92)
- Filters backslash hits by next character to avoid false positives
- Preserves the surrounding-character validation
- Returns immediately on first valid candidate
- Lets the tokenizer handle full validation (it already does this)

This reduces start() from O(n * delimiters * retries) to O(n) with
a very small constant factor per call.

* Update katex-extension.ts
2026-03-05 16:02:00 -06:00
Timothy Jaeryang Baek 700349064d chore: format 2026-01-08 01:55:56 +04:00
Timothy Jaeryang Baek c0ec04935b refac: citation 2025-12-26 02:05:03 +04:00
Timothy Jaeryang Baek ec45d77ce9 refac: sources and citations 2025-11-23 18:27:57 -05:00
Timothy Jaeryang Baek b0491886bc refac: disable single tilde 2025-11-23 17:16:10 -05:00
Tim Jaeryang Baek 9f9f1a1517 Merge pull request #17496 from ShirasawaSama/patch-16
feat: Dynamically load katex to improve first-screen loading speed (-630KB)
2025-09-17 01:36:55 -05:00
Timothy Jaeryang Baek c1f37d9aed refac 2025-09-17 01:22:15 -05:00
Shirasawa 9b3d71f0d2 feat: Dynamically load katex to improve first-screen loading speed 2025-09-17 04:18:16 +00:00
Timothy Jaeryang Baek 098f34f400 refac/enh: mention token rendering 2025-09-14 18:49:01 -04:00
silentoplayz 7a80e60785 fix(frontend): Attempt to resolve TypeError in RichTextInput.svelte
Fixes an issue where `ue.getWordAtDocPos is not a function` would be thrown in `MessageInput.svelte`.

The error was caused by a timing issue where the `getWordAtDocPos` method on the `RichTextInput` component was not available when called from an event handler within the same component.

This change refactors the code to pass the `getWordAtDocPos` function as a callback prop from `RichTextInput` to `MessageInput`, ensuring it's available when needed.
2025-08-03 22:36:08 -04:00
YuQX ec97fdf5d6 Fix: Solve equation problems with Chinese symbols by adding more symbols to the ALLOWED_SURROUNDING_CHARS 2025-05-23 15:06:37 +08:00
tth37 8e14372dd8 enh: Support more languages 2025-05-06 21:26:09 +08:00
tth37 de182ddec2 fix(katex): Allow Chinese characters adjacent to math delimiters 2025-05-06 18:38:42 +08:00
Timothy Jaeryang Baek d6f3444141 fix: katex issue 2025-02-19 23:37:11 -08:00
Timothy Jaeryang Baek 9feed97f22 refac: think tag 2025-01-22 09:24:40 -08:00
Timothy Jaeryang Baek c9dc7299c5 enh: <think> tag support 2025-01-22 00:13:24 -08:00
Yuta Hayashibe 12516c8a45 fix: Fix typos 2024-10-14 16:22:07 +09:00
Timothy J. Baek ffd598c5d7 enh: summary tag support 2024-09-30 12:50:53 +02:00
Timothy J. Baek 119a7f1933 doc: changelog 2024-09-25 15:45:36 +02:00
Hwang In Tak d501ece247 fix: Fix KaTeX corner cases 2024-09-25 02:15:53 +09:00
Timothy J. Baek e703e172e2 chore: format 2024-09-24 18:10:14 +02:00
Hwang In Tak 30e65b33f6 fix: Add comments 2024-09-25 00:41:08 +09:00
Hwang In Tak 3f1255b39e fix: Change inline and block delimiters 2024-09-25 00:10:49 +09:00
Hwang In Tak e48d66f918 fix: Remove unnecessary logging 2024-09-24 22:11:05 +09:00
Hwang In Tak 0bfbace9aa fix: Simplify regex 2024-09-24 22:00:01 +09:00
Hwang In Tak 377cc427b6 fix: Remove unnecessary logging 2024-09-24 20:40:50 +09:00
Hwang In Tak 214546399a fix: fix katex rendering 2024-09-24 16:58:15 +09:00
Timothy J. Baek 4f47053e93 refac 2024-08-16 17:51:50 +02:00
Timothy J. Baek 4ef042e966 refac 2024-08-16 15:33:14 +02:00