There is no FK between personas and persona_chat_conversations, and the
chat list skips conversations whose persona is gone — so deleting a persona
left its transcripts in the database, invisible and unreachable, forever.
Both deletes now run in one transaction so the two tables cannot get out of
step.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Persona chat stored a flat Vec<ChatMessage> keyed on (user_id, persona_id),
which meant one rolling transcript per persona and no way to revisit a turn.
This moves it onto the same ChatHistoryStore tree the file chat uses and
gives conversations their own identity.
Storage
- messages_json now holds a serialized ChatHistoryStore. Reads accept the
old flat array and upgrade it in place, so existing transcripts survive
without a data migration. The upgrade drops the v1 seed greeting, which
the flat renderer hid but the tree renderer would surface as a bubble the
user has never seen.
- New migration re-keys persona_chat_conversations on an opaque
conversation_id and adds title + created_at, so one persona can hold any
number of separate threads. Every DAO read and write is scoped by user_id
as well: a conversation id is a bearer token for someone's transcript and
must never grant access on its own.
- The per-conversation lock and in-flight turn slot key on conversation_id,
so two threads with the same persona can run turns concurrently.
Endpoints
- POST/DELETE /persona_chat/conversations — start and remove a thread,
replacing /persona_chat/reset.
- GET /persona_chat/conversations — the chat list, with snippet and counts
derived from each tree's active branch.
- POST /persona_chat/rewind, POST /persona_chat/switch-branch,
GET /persona_chat/branches — rewind and fork, mirroring the file chat.
Index 0 is rewindable here (it is the user's own first question, not a
synthetic prompt) and re-anchors on the seed node.
- history/turn/rewind/switch-branch/branches all key on conversation_id;
history gained branch_id and now returns real fork_info, active_leaf_id
and viewing_branch_id instead of placeholders.
The turn body no longer carries a persona at all — it is read from the
stored conversation, so a stale client cannot swap a thread's voice midway.
Titles
After the first turn persists, the conversation is named from its opening
exchange on the same backend the turn ran on. Small models wrap titles in
quotes, prefix them with "Title:" and append explanations, so sanitize_title
strips all of that and truncates on a character boundary. Any failure falls
back to the user's opening question. Generation runs after persistence: a
failed title must not cost the turn.
Fixes found along the way
- insight_chat: both file-chat turn paths captured path.len() before
apply_context_budget drained messages out of the middle, then sliced
messages[path_len..] for the new tree nodes. Once truncation fired that
dropped the user turn from the tree or panicked on an out-of-range start
index. Now read after the budget pass as history_len.
- Persona chat had no context budget at all and hardcoded truncated: false
in the done frame, so a rolling transcript grew unbounded.
- A cancelled turn persisted a half-finished transcript and pushed a second
terminal frame; it now returns early like the file chat.
- The seeded system prompt was frozen at conversation creation, so editing a
persona never reached a thread already in flight. Re-resolved per turn.
- turn_count was overwritten each write with the per-turn message delta;
it is now the cumulative assistant-turn count on the active branch.
- is_initial is always false: the file chat reserves it for its synthetic
"describe this photo" prompt, and marking a persona chat's first question
with it made the opening reply impossible to regenerate.
600 lib tests pass, clippy --all-targets clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The mobile client dispatches POST /persona_chat/turn with snake_case
keys (persona_id, user_message, num_ctx, ...) per the file-chat
convention, but the persona request/response structs carried camelCase
serde renames (personaId, userMessage, ...), so every dispatch 400'd
with 'missing field personaId'. Drop the renames from
PersonaChatTurnRequest, PersonaChatHistoryView, RenderedPersonaMessage,
and PersonaChatResetRequest — the persona wire shapes now match the
rest of the API (history query, 202 turn_id, SSE skip_before) — and
pin the contract with serialization round-trip tests.
- PersonaChatSession dispatches turns through the shared agent loop,
persisting transcripts keyed on (user_id, persona_id) in the new
persona_chat table (migration included).
- GET /persona_chat/history returns a rendered transcript (tool
invocations folded, is_initial flag) matching the file-chat shape.
- POST /persona_chat/turn returns 202 with a turn_id; SSE replay and
cancel reuse turn_replay_impl/cancel_turn_impl, extracted from the
actix-attributed handlers so both route families share the logic.
- POST /persona_chat/reset clears the persona transcript.
- Concurrent dispatches for the same (user, persona) are rejected with
409 via an in-flight gate (InFlightPersonaTurns) whose RAII guard
drops with the spawned turn task, freeing the slot on completion,
error, or abort.
HEIC/HEIF sources use Display P3 color primaries. Without
colorspace=bt709 the mjpeg encoder treated P3 values as sRGB,
producing warm/oversaturated output. Also bake EXIF Orientation
tag into pixels so saved JPEGs are canonically oriented.
- Extract orientation from exif-reader and pass through all
ffmpeg thumbnail paths (small, large, xlarge previews)
- Add shared build_image_thumb_filter() for the 200px path
- Add rotation + colorspace=bt709 to large/xlarge ffmpeg paths
Changed last_fork.clone() to last_fork.take() in render_tree_path
for both assistant and user branches. This prevents fork_info from
propagating to every downstream message, ensuring the chip appears
only at the actual divergence point in the conversation tree.
Rewind & Regenerate resends the identical question, so both branches'
first child matched and every picker option read the same. Snippets now
walk the branches' visible messages (tool scaffolding skipped) in
lockstep and preview the first position where the contents diverge; a
branch that runs out beforehand falls back to its own last message.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ForkInfo now carries the divergence node's id, and the branches endpoint
accepts node_id + viewing_branch_id to return only the position-ranked
sibling branches at that fork (snippets taken from after the divergence,
skipping tool scaffolding) instead of every leaf in the tree. Options in
the same subtree as the viewing leaf anchor to that leaf so clients can
reliably identify "the branch I'm on". Tree-wide listing is unchanged
when node_id is absent.
- llm_client: descendant_leaves + branch_options_at helpers; optional
position rank on BranchLeafInfo.
- insight_chat: ForkInfo.node_id set during render_tree_path; get_branches
takes the scoping params.
- handlers: dedicated ChatBranchesQuery; unknown node_id maps to 400.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Store chat history as a tree in training_messages instead of a flat
array. Rewind now sets active_leaf_id to the target node rather than
truncating, so discarded paths survive as alternate branches. Fork
indicators ("X/Y") mark divergence points and let the client load or
switch to alternate paths.
- llm_client: StoredChatNode/ChatHistoryStore with tree traversal
helpers (path_to_leaf, children_of, fork_at_node, leaves*, etc.) and
from_flat_array for backward-compatible reads of the old flat format.
- insight_chat: load_history/chat_turn/rewind_history operate on the
tree; new switch_branch and get_branches. render_tree_path walks every
node (including non-rendered tool-dispatch nodes) when computing fork
info, so a fork whose diverging child is a tool call still surfaces on
the following rendered message.
- handlers: GET /insights/chat/branches, POST /insights/chat/switch-branch,
and a branch_id param on the history endpoint.
Backward-compatible: flat arrays are converted to a tree lazily on read.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add retry_with_backoff utility (100ms base, ±25% jitter, 3 attempts)
and replace all 11 direct DAO calls in insight_chat.rs and
insight_generator.rs. The Mutex guard drops between retries so the
maintenance connection can release its SQLite write lock.
New optional SamplingOverride forwarded to llama-server as
chat_template_kwargs.enable_thinking (gates Qwen3-style reasoning
blocks). None leaves the template default; other backends ignore it.
Wired through the agentic-insight and chat-turn request bodies/handlers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ALL-mode over-constrains NL queries — the model maps several query words to
tags and few photos carry every one, zeroing the candidate set. Switch to
ANY (a photo matches if it has any named tag); the semantic CLIP ranking
provides precision within that pool. Exclude tags still filter out.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When structured filters are present they're the constraint and CLIP only ranks
within the candidate set, so drop the global similarity threshold for that
case. Previously the 0.2 whole-library threshold ran BEFORE intersecting with
the filters, discarding filter-matching photos that scored just under it (e.g.
a 2022 beach photo at 0.18) — producing after_struct_filter=0 even when matches
existed. Plain semantic (no filters) keeps the user's threshold.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The CLIP encode failure reason was only ever returned in the HTTP response
body, never logged server-side, making 502s from /photos/search opaque. Log
the underlying cause — network error to the URL, or the Apollo HTTP status +
response body — so CLIP-service problems are diagnosable from the ImageApi log.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pin the NL->structured translation to a small, fast model that can stay
co-resident with CLIP (and the chat model) so it never evicts them on a tight
VRAM budget. Precedence: UNIFIED_SEARCH_MODEL env > client-selected model >
configured default. Logs the effective model (backend.model()) so model A/B
tests are visible. Documented in .env.example.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Log the translated query (semantic/tags/place/date/media + has_struct), the
tag-filter file count, candidate-row + allowed-hash counts, and the CLIP
considered/hits/after-filter counts. Pinpoints which stage drops results to
zero (over-extracted filter, tag path mismatch, Any/All over-constraint, or
CLIP threshold). info-level for now while debugging.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add an optional `model` query param to /photos/search/unified, passed into
resolve_backend's overrides. The client sends the user's currently-selected
local model so the translation step reuses an already-loaded model instead of
forcing a llama-swap eviction + cold start. Falls back to the configured
default when absent. Still local only (no hybrid).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Composes the two existing engines (Path A orchestration):
- Translate NL -> StructuredQuery via local LLM, respecting LLM_BACKEND
(resolve_backend(Local) -> ollama or llama-swap; no hybrid).
- Forward-geocode the place name into a gps circle.
- Structured filters (tags/EXIF/geo/date/media) build a candidate set of EXIF
rows; CLIP ranks within it, joined by content_hash. Degenerate cases match
existing behavior: semantic-only -> plain CLIP; filters-only -> date-sorted.
- Echoes the interpreted query (incl. resolved place) for editable client chips.
Refactor: extracted reusable cores from clip_search (score_photos, resolve_hits,
parse_library_scope, score_error_response) shared by both endpoints. Removed the
Phase 1 allow-until-wired attributes now that nl_query + geo are consumed.
fmt + clippy clean; 23 backend tests pass (7 geo, 12 nl_query, 4 unified).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Foundation for the /photos/search/unified endpoint (Phase 2). Two new,
fully unit-tested pieces, not yet wired into a route (allow-until-wired,
mirroring llm_client.rs):
- ai/nl_query.rs: translate a free-text query into a StructuredQuery via one
grounded LLM call. Two-stage — the model emits names/ISO dates, then a pure
resolve step maps tag names against the real vocab and converts dates to
unix seconds. Hallucinated (non-vocab) tags are surfaced in unmatched_tags
rather than silently used as hard filters — the anti-noise guard. 12 tests.
- geo::forward_geocode + bbox_to_circle: resolve a place name to a circle via
Nominatim /search, collapsing the bounding box to centroid + circumscribing
radius so "Portland" and "Italy" both map onto the existing gps circle
filter with no schema change. Radius is the max centroid-to-corner distance
(corners aren't equidistant on a sphere). 4 tests.
fmt + clippy clean; 19 new tests pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Nothing reaped reels before, so the on-disk cache and ledger grew
unbounded — each night's daily reel is a new ~4MB file + ledger row that's
stale within ~26h.
- Pre-gen self-prune: after recording a reel, prune_superseded keeps the
newest PREGEN_KEEP_PER_SCOPE (2) rows per (span, library) and unlinks the
superseded reels' mp4+sidecar. Caps the ledger/disk at ~spans×libraries×2.
- On-disk sweeper (spawn_reel_cache_sweeper): every 24h, removes reel mp4s
with no ledger row and no live job older than REEL_CACHE_MAX_AGE_DAYS (7) —
bounding the on-demand cache, which has no ledger row and otherwise grows
forever — plus crashed-render cruft (.mp4.tmp/.concat.txt/orphan sidecars).
Runs regardless of REEL_PREGEN_ENABLED; disable with REEL_CACHE_SWEEP_ENABLED=0.
- New DAO methods prune_superseded + all_cache_keys (with tests); env knobs
documented in .env.example.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Past the key-aware dedup, any mp4 already at the cache key was not
pre-generated by us (no matching ledger row) — typically an on-demand
fast-scripted reel sharing the key after the max_segments alignment.
Adopting it recorded a ledger row pointing at the fast reel, silently
defeating agentic pre-gen. Drop the adopt-existing-mp4 shortcut and
always produce_reel (atomic overwrite). Worst case is one redundant
re-render if a prior run crashed between render and ledger write.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
exists_fresh only matched (span, library, render_version, age), so a
cache-key change that doesn't bump RENDER_VERSION (e.g. the max_segments
alignment, or any future selection-logic tweak) left last night's ledger
row looking 'fresh' — the nightly run would skip and the orphaned reel
would persist. Dedup now compares the stored cache_key to the freshly
computed key (and confirms the mp4 exists), so a changed key forces a
regen within the freshness window. exists_fresh stays as the HTTP
endpoint's fast gate.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
pregen_one hardcoded max_segments: 24 while create_reel_handler defaults
to DEFAULT_MAX_SEGMENTS (40). Since the cache key encodes the raw
max_segments, the pre-generated reel's key never matched the client's
on-demand request, so POST /reels cache-hit an older max=40 reel and the
agentic pre-gen file was left orphaned. Align to DEFAULT_MAX_SEGMENTS (as
the plan specified) so the on-demand cache-hit path serves the pre-gen
reel. Content is unchanged — the actual beat count is duration-budgeted
either way; only the key descriptor differed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- pregen_one recorded media_count as planned.len() (beat count); record
the actual media item total (media.len(), photos + clips) in both the
cache-hit and freshly-rendered ledger paths. Drops the redundant
photo_count binding.
- Replace upsert_prefs's insert-then-catch-error-then-update dance with a
single atomic INSERT ... ON CONFLICT(id) DO UPDATE. Explicit id=1 makes
the conflict target deterministic; explicit column .set((...)) keeps
None -> NULL overwrite semantics so the row mirrors the latest request
exactly, and genuine insert errors surface instead of being swallowed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1. Drop the unregistered prefs_dao/reel_dao web::Data extractors from
create_reel_handler / precomputed_reel_handler and read the DAOs off
AppState instead (consistent with the scheduler). Missing app_data
would have 500'd every POST /reels and /reels/precomputed at runtime.
2. Restore the dropped 'return' in the cache-hit branch — without it a
cache hit fell through, overwrote the Done job with Queued, and
re-ran the whole TTS+render pipeline on every request.
3. Make secs_until_next_run_hour minute/second-accurate so a batch that
finishes inside the run hour sleeps ~24h instead of busy-looping
(wake, re-run, sleep 0) for the rest of the hour. Tests updated.
4. Prune photo/user-bound tools (get_file_tags, get_faces_in_photo,
recall_facts_for_photo, recall_facts_for_entity) from the agentic
reel scripter's allow-list — they no-op/error with the empty
file/user context and only burn iterations.
5. Align AGENTIC_SYSTEM_PROMPT's advertised tool list with the actual
(pruned) allow-list.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Implement end-to-end nightly pre-generation of memory reels with agentic
scripting that grounds narration in calendar, location, messages, and RAG.
Sections A-E from the plan:
A. Extract produce_reel pipeline core from run_reel_job with
ScripterMode::Fast/Agentic and progress callbacks.
B. Agentic scripter: factor run_readonly_tool_loop from the insight
generator, build read-only tool gate, prompt builder with GPS, and
generate_script_agentic with fallback to fast path.
C. Precomputed reels ledger (SQLite table + DAO), GET /reels/precomputed
handler with validity gate, GET /reels/by-key/{key}/video streaming,
and normalize_library_key helper.
D. Nightly scheduler: spawn_pregen_scheduler with configurable hour,
run_pregen_batch (day/week/month spans), pregen_one with dedup and
disk-check, secs_until_next_run_hour time math.
E. user_ai_prefs passive mirror table + DAO for param capture in
create_reel_handler and replay in the scheduler.
Also fixes resolve_library_param signature to take &[Library] and adds
resolve_library_param_state wrapper for AppState callers.
New files: migrations/2026-06-13-000000_add_precomputed_reels/,
migrations/2026-06-13-000010_add_user_ai_prefs/,
src/database/precomputed_reel_dao.rs,
src/database/user_ai_prefs_dao.rs
A clip beat capped playback at CLIP_SECONDS and filled the rest of the
narration with a tpad freeze-frame, so a clip stopped dead on its last
frame for a second or two before the transition — a glitchy pause that
stills don't have. Extract clip_beat_plan: the clip now plays for as
much of its beat as the source footage covers, and we freeze only when
the source is genuinely shorter than the narration. Bump RENDER_VERSION.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
0.08s read as too abrupt; 0.12s keeps the burst clearly snappier than the
0.35s held-shot fade without jarring. Bumps RENDER_VERSION.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Videos in a span now appear as clip beats: the first few seconds of the
video (capped at CLIP_SECONDS=5, and to the source length) filled to the
portrait canvas like photos, with its live audio ducked under the
narration (amix at 0.35). If the narration outlasts the clip, the last
frame is held (tpad); clips with no audio track just play under narration.
Selection splits the beat budget between photo beats and clip beats —
clips get up to half (≥1 when present), photos the rest — then merges
both back into chronological order. SegmentMedia gains a Clip variant;
beats carry `media` (photos or one clip) and the cache key tags P/C so a
path used as a still vs a clip differ.
Also drops the burst fade from 0.15s to 0.08s so a quick burst reads
clearly differently from a held shot. Bumps RENDER_VERSION.
The clip filtergraph (fill + duck-mix + last-frame hold) is unit-tested
but, like the rest of the ffmpeg path, wants a real render check on the
GPU host.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Restructures a reel around beats — one narration line over one or more
photos — instead of one line per photo. A single-photo beat is a held
shot; a multi-photo beat is a quick burst that flashes through several
moments of an event while the line is read. So a week/month reel can show
everything it spans without a narrated (and timed) segment per photo.
Selection (selector.rs):
- Duration budget: cap the number of narrated beats to ~REEL_TARGET_SECONDS
(default 90, env-tunable) so week/month reels don't run minutes long.
- Event clustering by time gap; when there are more events than the beat
budget, adjacent events merge so the whole span stays covered. Each beat
bursts up to MAX_BURST_PHOTOS (an even spread), so a 40-shot dinner
contributes a handful of quick frames, not forty narrated seconds.
Render (render.rs): a beat renders its photos as a concat of per-photo
fills (blurred-bg portrait, fps-before-fade) under one muxed narration;
burst photos get a snappier fade. beat_durations splits the narration
across the photos, stretching only if a long burst would flash too fast.
Adds high-level info logs across the steps (request → script → per-beat
narrate/render → join → done with elapsed) for visibility. Bumps
RENDER_VERSION to re-render cached reels.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The fade looked steppy/low-frame-rate because the filtergraph normalized
fps AFTER the fade filters: the brightness ramp was sampled at the looped
still's coarse input cadence, then duplicated up to 30fps. Move fps ahead
of the fades, pin the still's input framerate (-framerate), and force CFR
output (-r) so the dip ramps across a full 30 frames and plays steadily.
Ease narration expressiveness from 0.7 to 0.6 (still tunable via
REEL_TTS_EXAGGERATION). Bump RENDER_VERSION so existing reels re-render.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fixes the "image is tiny" problem: a 1920x1080 landscape reel letterboxes
to a ~25%-height band on a portrait phone. Switch to a portrait 1080x1920
canvas and fill it per photo with a blurred, zoomed copy of the image
behind the sharp fitted photo — so the frame is always full regardless of
the photo's orientation, with no black bars and no cropping of the subject.
Add a quick 0.35s fade in/out baked into each segment so concatenated
photos dip smoothly instead of hard-cutting (fade-out lands in the
narration's silent tail, so speech isn't clipped). Drop the unused
Ken Burns branch — motion can return deliberately later.
Warm up the narration a touch: thread Chatterbox's `exaggeration` through
synthesize_serialized and default reels to 0.7 (tunable via
REEL_TTS_EXAGGERATION). Bump RENDER_VERSION so existing landscape reels
re-render.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The concat stage wrote to <key>.mp4.tmp (for an atomic publish-rename),
but ffmpeg infers the muxer from the output extension and can't map
.tmp to a format — "Unable to choose an output format". Force the mp4
muxer explicitly so the temp extension is irrelevant. Segment render,
NVENC, TTS, and scripting were already working end-to-end; this was the
only failure, at the final join.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
New POST /reels + GET /reels/{id} (+ /video) build an MP4 slideshow of a
memory span (day/week/month), narrated by the LLM in a cloned voice.
Pipeline (src/reels/): a selector resolves which photos + reel metadata,
the scripter writes one narration line per photo via a single LLM call
(reusing each photo's cached insight as context — no fresh vision calls,
so reel generation stays off the GPU's vision slot), each line is
synthesized to speech, and the renderer assembles stills + narration via
ffmpeg. Jobs run in the background (mirroring the TTS speech-job
registry) since a reel takes minutes; the finished MP4 is cached on disk
keyed by the selection so a repeat request is instant.
The segment model is media-typed (Photo today) so a video-clip segment
(phase 2) and a nightly pre-render (phase 3) slot in without reworking
the pipeline. Ken Burns motion is implemented but defaulted off pending a
visual check on the GPU box.
Supporting changes:
- memories: extract gather_memory_items() so the reel selector reuses the
exact window/exclusion/tz/sort logic behind /memories.
- ai::tts: add synthesize_serialized() so reel narration honors the same
single-GPU permit + write lease as user TTS requests.
- video::ffmpeg: make get_duration_seconds() pub for narration timing.
- AppState: reels_path (REELS_DIRECTORY, defaults beside preview clips).
Pure logic (cache key, script parsing, ffmpeg arg/filter construction,
even sampling, segment timing) is unit-tested (26 tests). The runtime
path (ffmpeg render, TTS, LLM) needs a real run on the GPU host to verify
end-to-end — not exercisable in CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Clones that don't start at 0:00 are tagged with where the reference
window begins (grandma-at1m32s-30s), so voices cloned from different
sections of the same source are distinguishable in the voice list.
Zero-start names keep the existing -30s form.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both voice creation endpoints (upload + from-library) now accept optional
start_seconds/duration_seconds, threaded to ffmpeg as -ss/-t, so the
reference window can target clean speech anywhere in a long recording
instead of always the first N seconds. Duration is clamped to the
LLAMA_SWAP_TTS_REF_SECONDS cap and the voice-name tag reflects the
actual window length.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A JSON map (TTS_PRONUNCIATIONS_PATH, default tts_pronunciations.json)
rewrites mispronounced words — place names, initialisms, dotted
abbreviations — to phonetic spellings before synthesis, applied after
markdown cleanup in both /tts/speech paths. Whole-word smartcase
matching (lowercase keys match any casing, uppercase keys exact),
longest key wins, hot-reloaded on mtime change with last-good fallback
on parse errors. See tts_pronunciations.example.json.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Models wrap the title line despite the prompt — "**Title: A Day in the
Woods**", "## Title: ...", bold around just the label — which made
parse_title_body's bare "Title:" prefix match fall through to the
fallbacks and leak asterisks into the stored title.
strip_title_markdown trims bold/italic markers, heading hashes,
backticks, and quotes from both ends; applied to the label line, the
extracted title, both fallback paths, and generate_photo_title (which
previously stripped only quotes).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Trialing Qwen3-Embedding-0.6B (1024-dim, instruct-prefixed queries)
against nomic required code changes at every hardcoded seam; now it's a
config flip plus a reembed_embeddings run.
- EMBEDDING_DIM env (default 768) replaces every hardcoded dim check:
daily summary / calendar / search / location DAOs, Ollama batch
validation, reembed_embeddings
- entities gains the dim guard it never had — a wrong-dim vector
silently kills dedup/recall (cosine over mismatched lengths is 0),
so store None and warn instead
- embed_query / embed_document split with EMBED_QUERY_PREFIX /
EMBED_DOCUMENT_PREFIX (literal \n expanded): retrieval models treat
the two sides differently — nomic wants search_query:/search_document:,
Qwen3 wants Instruct:...\nQuery: on queries only. All query-side
call sites and all corpus writers now declare their side.
- document the contract in CLAUDE.md: change the model or any of these
vars → re-run reembed_embeddings or search is garbage
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The GPU lease keeps per-request reqwest budgets from burning behind a
cross-model swap, but the job-level INSIGHT_GENERATION_TIMEOUT_SECS
wall-clock started at spawn — an insight queued behind a running TTS
synthesis parked its first chat call on the lease and timed out
("timeout after 180s") before chatterbox even finished loading.
Acquire-and-drop an LLM read lease before starting the job clock in
both insight handlers: the wait for the GPU happens before the
timeout begins, mirroring the per-request lease semantics. Dropped
immediately — holding it across the generation would deadlock the
chat calls' own lease acquisitions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Queries embedded via llama-swap were searching corpora embedded via
Ollama (measured: spaces diverged). Introduce LocalLlm — the local
Ollama + llama-swap pair with LLM_BACKEND dispatch baked in — and route
all embedding writers through it; anything embedding via a concrete
client reintroduces the bug.
- search_rag: embed the model's query verbatim (no metadata boilerplate),
make date optional — no time-decay when omitted, so "when did X
happen?" queries rank purely by similarity across all time
- reembed_embeddings bin: re-embed summaries / calendar / search /
knowledge entities via the active backend, with old-new cosine report
per table and truncate-and-retry for inputs over the embed server's
physical batch size
- import_calendar, import_search_history: embed through LocalLlm
- search_messages / get_sms_messages: render sender → recipient so sent
messages are attributable to a conversation
- insight job failures: store the one-line anyhow context chain ({:#})
instead of the Debug dump the client was shown verbatim
- serialize env_dispatch tests behind a lock (parallel-runner flake)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>