T

Cameron 0aaea91cc2 feat: add content_hash backfill + register every media file

Adds blake3 content hashing as the basis for derivative dedup
(thumbnails, HLS) across libraries. Computed inline by the watcher on
ingest and by a new `backfill_hashes` binary for historical rows.

Key changes:
- `content_hash` and `size_bytes` are now populated on new image_exif
  rows; a new ExifDao surface (`get_rows_missing_hash`,
  `backfill_content_hash`, `find_by_content_hash`) supports backfill and
  future hash-keyed lookups.
- The watcher now registers every image/video in image_exif, not just
  files with parseable EXIF. EXIF becomes optional enrichment; videos
  and other non-EXIF files still get a hashed row. This also makes
  DB-indexed sort/filter cover the full library.
- `/image` thumbnail serve dual-looks up hash-keyed path first, then
  falls back to the legacy mirrored layout.
- Upload flow accepts `?library=` query param + hashes uploaded files.
- Store_exif logs the underlying Diesel error on insert failure so
  constraint violations surface instead of hiding behind a generic
  InsertError.
- New migration normalizes rel_path separators to forward slash across
  all tables, deduplicating any rows that collide after normalization.
  Fixes spurious UNIQUE violations from mixed backslash/forward-slash
  paths on Windows ingest.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

2026-04-21 01:55:07 +00:00

.claude/commands

Add Speckit and Constitution

2026-02-26 10:05:47 -05:00

.idea

Build insight title from generated summary

2026-02-24 16:08:25 -05:00

.specify

Add Speckit and Constitution

2026-02-26 10:05:47 -05:00

migrations

feat: add content_hash backfill + register every media file

2026-04-21 01:55:07 +00:00

specs/001-video-wall

Add VideoWall feature: server-side preview clip generation and mobile grid view

2026-02-25 19:40:17 -05:00

src

feat: add content_hash backfill + register every media file

2026-04-21 01:55:07 +00:00

.gitignore

Create Insight Generation Feature

2026-01-03 10:30:37 -05:00

Cargo.lock

feat: add content_hash backfill + register every media file

2026-04-21 01:55:07 +00:00

Cargo.toml

feat: add content_hash backfill + register every media file

2026-04-21 01:55:07 +00:00

CLAUDE.md

Add comprehensive testing for preview clip and status handling

2026-02-26 10:06:21 -05:00

diesel.toml

Move database into the main app

2020-07-07 21:48:29 -04:00

Jenkinsfile

Update CI to Rust 1.59

2022-03-01 20:44:51 -05:00

README.md

feat: add model-availability validation to agentic insight generation (T009-T011)

2026-03-18 23:07:43 -04:00

README.md

Image API

This is an Actix-web server for serving images and videos from a filesystem. Upon first run it will generate thumbnails for all images and videos at BASE_PATH.

Features

Automatic thumbnail generation for images and videos
EXIF data extraction and storage for photos
File watching with NFS support (polling-based)
Video streaming with HLS
Tag-based organization
Memories API for browsing photos by date
Video Wall - Auto-generated short preview clips for videos, served via a grid view
AI-Powered Photo Insights - Generate contextual insights from photos using LLMs
RAG-based Context Retrieval - Semantic search over daily conversation summaries
Automatic Daily Summaries - LLM-generated summaries of daily conversations with embeddings

Environment

There are a handful of required environment variables to have the API run. They should be defined where the binary is located or above it in an .env file. You must have ffmpeg installed for streaming video and generating video thumbnails.

DATABASE_URL is a path or url to a database (currently only SQLite is tested)
BASE_PATH is the root from which you want to serve images and videos
THUMBNAILS is a path where generated thumbnails should be stored
VIDEO_PATH is a path where HLS playlists and video parts should be stored
GIFS_DIRECTORY is a path where generated video GIF thumbnails should be stored
BIND_URL is the url and port to bind to (typically your own IP address)
SECRET_KEY is the hopefully random string to sign Tokens with
RUST_LOG is one of off, error, warn, info, debug, trace, from least to most noisy [error is default]
EXCLUDED_DIRS is a comma separated list of directories to exclude from the Memories API
PREVIEW_CLIPS_DIRECTORY (optional) is a path where generated video preview clips should be stored [default: preview_clips]
WATCH_QUICK_INTERVAL_SECONDS (optional) is the interval in seconds for quick file scans [default: 60]
WATCH_FULL_INTERVAL_SECONDS (optional) is the interval in seconds for full file scans [default: 3600]

AI Insights Configuration (Optional)

The following environment variables configure AI-powered photo insights and daily conversation summaries:

Ollama Configuration

OLLAMA_PRIMARY_URL - Primary Ollama server URL [default: http://localhost:11434]
- Example: http://desktop:11434 (your main/powerful server)
OLLAMA_FALLBACK_URL - Fallback Ollama server URL (optional)
- Example: http://server:11434 (always-on backup server)
OLLAMA_PRIMARY_MODEL - Model to use on primary server [default: nemotron-3-nano:30b]
- Example: nemotron-3-nano:30b, llama3.2:3b, etc.
OLLAMA_FALLBACK_MODEL - Model to use on fallback server (optional)
- If not set, uses OLLAMA_PRIMARY_MODEL on fallback server

Legacy Variables (still supported):

OLLAMA_URL - Used if OLLAMA_PRIMARY_URL not set
OLLAMA_MODEL - Used if OLLAMA_PRIMARY_MODEL not set

SMS API Configuration

SMS_API_URL - URL to SMS message API [default: http://localhost:8000]
- Used to fetch conversation data for context in insights
SMS_API_TOKEN - Authentication token for SMS API (optional)

Agentic Insight Generation

AGENTIC_MAX_ITERATIONS - Maximum tool-call iterations per agentic insight request [default: 10]
- Controls how many times the model can invoke tools before being forced to produce a final answer
- Increase for more thorough context gathering; decrease to limit response time

Fallback Behavior

Primary server is tried first with 5-second connection timeout
On failure, automatically falls back to secondary server (if configured)
Total request timeout is 120 seconds to accommodate LLM inference
Logs indicate which server/model was used and any failover attempts

Daily Summary Generation

Daily conversation summaries are generated automatically on server startup. Configure in src/main.rs:

Date range for summary generation
Contacts to process
Model version used for embeddings: nomic-embed-text:v1.5