Commit Graph
1136 Commits
Author SHA1 Message Date
Timothy Jaeryang Baek c653e4ec54 refac 2026-02-12 15:25:24 -06:00
Classic298andClaude 8cf32ae2a7 fix: prevent worker death during document upload by using run_coroutine_threadsafe (#21158)
* fix: prevent worker death during document upload by using run_coroutine_threadsafe

Replace asyncio.run() with asyncio.run_coroutine_threadsafe() in
save_docs_to_vector_db() to prevent uvicorn worker health check failures.

The issue: asyncio.run() creates a new event loop and blocks the thread
completely, preventing the worker from responding to health checks during
long-running embedding operations (>5 seconds default timeout).

The fix: Schedule the async embedding work on the main event loop using
run_coroutine_threadsafe(). This keeps the main loop responsive to health
check pings while the sync caller waits for the result.

Changes:
- main.py: Store main event loop reference in app.state.main_loop at startup
- retrieval.py: Use run_coroutine_threadsafe() instead of asyncio.run()

https://claude.ai/code/session_01UQSYvSTkXb57sFb7M85Kcw

* add env var

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-02-12 15:22:57 -06:00
Timothy Jaeryang Baek 531ac70ce5 refac 2026-02-12 12:07:45 -06:00
Timothy Jaeryang Baek 423d8b1817 refac 2026-02-12 11:04:34 -06:00
Timothy Jaeryang Baek 05ae44b98d refac 2026-02-12 11:01:46 -06:00
Timothy Jaeryang Baek a40808579f refac 2026-02-12 10:59:41 -06:00
Timothy Jaeryang Baek ccb71a7322 refac 2026-02-11 18:32:14 -06:00
Classic298andMichael efe5416f83 fix: reduce TTFT by caching model lookups in chat completion (#20886)
fix: reduce TTFT by caching model lookups in chat completion

Skip expensive get_all_models() calls when models are already cached
in app.state. This significantly reduces Time To First Token (TTFT)
for chat completions and embeddings requests.

Previously, every request called get_all_models() which fetches model
lists from all configured backends. Now we check the cache first and
only call get_all_models() on cache miss.

Affected endpoints:
- openai: generate_chat_completion, embeddings
- ollama: embed, embeddings

Fixes #20069

Co-authored-by: Michael <42099345+mickeytheseal@users.noreply.github.com>
2026-02-11 18:29:10 -06:00
Timothy Jaeryang Baek a4281f6a7f refac: ldap 2026-02-11 18:25:37 -06:00
Timothy Jaeryang Baek 2372b70031 refac: async pipelines requests 2026-02-11 18:24:30 -06:00
Timothy Jaeryang BaekandClassic298 dddac2b0ca refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-02-11 18:19:01 -06:00
Timothy Jaeryang BaekandClassic298 0da57149ae refac
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
2026-02-11 18:13:30 -06:00
Thomas Rehn ce3a615442 perf: cache OpenAI config reads to avoid redundant Redis lookups in /api/models (#21306)
Each access to request.app.state.config.<KEY> triggers a synchronous
Redis GET. In get_all_models_responses() and get_merged_models(), the
config keys OPENAI_API_BASE_URLS, OPENAI_API_KEYS, and
OPENAI_API_CONFIGS were read on every loop iteration — resulting
in some cases in 200-300 Redis round-trips for OPENAI_API_BASE_URLS alone.

Read each config value once into a local variable at the start of the
function and reuse it throughout.
2026-02-11 17:59:50 -06:00
Timothy Jaeryang BaekandJannik S. 4e0cb88583 refac: audio timeout
Co-Authored-By: Jannik S. <jannik@streidl.dev>
2026-02-11 17:56:49 -06:00
Timothy Jaeryang Baek 9b30e8f689 refac 2026-02-11 17:53:01 -06:00
Timothy Jaeryang Baek f376d4f378 chore: format 2026-02-11 16:24:11 -06:00
Timothy Jaeryang Baek e5035ea31e refac 2026-02-11 15:55:23 -06:00
Timothy Jaeryang Baek c8cbdc8f7f refac 2026-02-11 15:24:12 -06:00
Timothy Jaeryang Baek 64c37ab968 refac 2026-02-11 15:12:37 -06:00
Timothy Jaeryang Baek a38ad8fc42 refac 2026-02-11 14:09:55 -06:00
Timothy Jaeryang Baek c2207887b3 feat: skills backend 2026-02-11 14:00:34 -06:00
Timothy Jaeryang Baek 3e56261c5e refac 2026-02-11 02:06:43 -06:00
Timothy Jaeryang Baek 30f72672fa refac 2026-02-10 15:57:08 -06:00
Timothy Jaeryang Baek 4aedfdc547 refac 2026-02-10 15:47:21 -06:00
Timothy Jaeryang Baek e3a8257690 refac 2026-02-10 15:41:11 -06:00
Timothy Jaeryang Baek c259c87806 refac 2026-02-10 15:30:16 -06:00
Classic298 f236192fe1 fix: resolve N+1 query in knowledge batch file add (#21006) 2026-02-09 16:17:30 -06:00
Timothy Jaeryang Baek c2f5cb542e refac 2026-02-09 14:03:35 -06:00
Tim Baek 48a0abb40f Merge pull request #21277 from open-webui/acl
refac: acl
2026-02-09 13:34:36 -06:00
Timothy Jaeryang Baek f7406ff576 refac 2026-02-09 13:28:14 -06:00
Tim Baek e2d09ac361 refac 2026-02-09 09:06:48 +04:00
Tim Baek 4852227158 refac 2026-02-08 06:22:56 +04:00
Classic298 494cf8b3ef fix (#21226) 2026-02-08 03:19:26 +04:00
Tim Baek 258454276e fix: files settings save issue 2026-02-06 22:33:49 +04:00
Tim Baek 2c37daef86 refac 2026-02-06 03:23:37 +04:00
G30 cac5dd12e9 fix: handle null data in model_response_handler (#21112)
Fix `AttributeError` in `model_response_handler` when processing channel messages with `null` data field. The function iterates over thread messages to build conversation history, but some messages may have `data=None` causing a crash when accessing `thread_message.data.get()`. Added null check using `(thread_message.data or {}).get("files", [])` to safely handle messages without data.
2026-02-05 15:15:34 -05:00
Timothy Jaeryang Baek e62649f940 enh: analytics 2026-02-05 00:00:49 -06:00
Timothy Jaeryang Baek 68a1e87b66 enh: analytics model modal 2026-02-04 23:42:46 -06:00
Timothy Jaeryang Baek ecf3fa2feb refac 2026-02-03 23:36:15 -06:00
Tim Baek cfd30581d5 Merge branch 'dev' into chat-message-rebased 2026-02-02 09:33:41 -06:00
Timothy Jaeryang Baek ea9c58ea80 feat: experimental responses api support 2026-02-01 19:39:28 -06:00
Tim Baek b2c2f1bd49 refac 2026-02-01 10:24:04 +04:00
Tim Baek 679e56c494 feat: token analytics 2026-02-01 10:19:59 +04:00
Tim Baek 599cd2eeeb feat: analytics backend API with chat_message table
- Add chat_message table for message-level analytics with usage JSON field
- Add migration to backfill from existing chats
- Add /analytics endpoints: summary, models, users, daily
- Support hourly/daily granularity for time-series data
- Fill missing days/hours in date range
2026-02-01 07:04:13 +04:00
Timothy Jaeryang BaekandHsienz f9ab66f51a refac
Co-Authored-By: Hsienz <55347238+hsienz@users.noreply.github.com>
2026-01-30 00:46:42 +04:00
Classic298 e686554392 fix: resolve N+1 query in SCIM group_to_scim user lookup (#21005) 2026-01-29 21:43:33 +04:00
Timothy Jaeryang Baek 93ed4ae2cd enh: files data controls 2026-01-29 19:50:06 +04:00
Timothy Jaeryang Baek a10ac774ab enh: manage shared chats 2026-01-29 18:51:02 +04:00
Timothy Jaeryang Baek ce50d9bac4 refac 2026-01-28 01:14:22 +04:00
7. Sun 33020d826f perf: parallelize image loading in image_edits endpoint (#20911)
Use asyncio.gather() to load multiple images concurrently instead of
sequentially, significantly reducing latency for multi-image edit
operations.
2026-01-28 00:35:25 +04:00