Legal
Subprocessors
Active subprocessors that process Sparky data. Sparky will update this list before a new subprocessor begins processing user data.
Active subprocessors
| Processor | Purpose | Data processed |
|---|---|---|
| Cloudflare R2 | Canonical document object storage and database backup target | Document files, object keys, hashes, database backup dumps |
| Cloudflare (network edge and Tunnel) | Delivers sparkyfetch.me to the internet, terminates TLS, and reports which crawlers fetched which pages | IP addresses, user agents, request paths, and response codes seen at the edge as traffic passes through. Sparky reads only aggregate counts from this — never per-visitor records. |
| Backblaze B2 | Cross-provider document archive | Mirrored document object copies |
| Resend | Magic-link email, invite email, operator notifications | Email addresses, email content, delivery metadata |
| Google (Sign in with Google) | Optional sign-in provider when used by a user | Email, profile name, avatar, OAuth account id |
| Google Drive (Drive API) | Optional one-way sync of documents to a user's own Google Drive when connected (drive.file scope) | Copies of the user's document files and folder names, plus an encrypted OAuth refresh token for the user's Google account |
| Telegram | Telegram bot commands, uploads, and replies when a user links Telegram | Telegram ids, usernames, messages, attachments |
| Langfuse (US Cloud) | LLM tracing, feedback, and cost observability | Trace metadata, prompts/outputs when content recording is enabled, user/session ids. Content is redacted through the L0–L3 layered stack when content recording and redaction are active. |
| Sentry (sentry.io, EU region) | Runtime error and crash monitoring for both the web and the document-processing service | Stack traces, URL paths, error messages, sanitized request headers, hashed account identifiers. Document text, OCR content, chat content, cookies, and raw IP addresses are stripped before transmission. |
| PostHog (eu.posthog.com, EU region) | Server-side, cookieless product analytics | Event names, salted SHA-256 hashes of account and household identifiers, plan/tier metadata, structured event properties. No browser SDK, no cookies, no session recording, no raw emails, filenames, or chat content. |
| OpenRouter | Routing layer for AI model calls to downstream model providers | Prompts, document text, metadata, chat context, model usage |
| Google (Gemini, direct) | Emergency embeddings fallback only, used if the primary OpenRouter embeddings route is unavailable. OCR and primary embeddings are routed through OpenRouter. | Document text chunks for embedding computation |
| OpenAI | Embeddings fallback via OpenAI API when configured | Document text chunks for embedding computation |
| ZeroEntropy | Memory reranking (reranker model) for the AI assistant's recall | Recalled memory-card text and the query used to re-rank them |
| Hindsight (self-hosted) | Self-hosted memory service. Internally calls third-party LLM providers (Groq/Gemini) for fact extraction from chat content and OpenRouter for embedding generation. Stores extracted facts and memory cards for the AI assistant's recall. | Chat content sent for extraction, memory observations, retained cards, recall metadata, source tags, document excerpts. Only email addresses are scrubbed before content leaves Hindsight for extraction or embedding. Phone numbers, addresses, document numbers, and other personal identifiers may be included in the extraction payload. |
| Cloudflare (Tunnel, DNS, edge) | HTTPS ingress, DNS, tunnel, and edge security | Request metadata, IP-derived logs, TLS and routing metadata |
Changes to this list
Sparky will update this list and the Privacy Policy before a new subprocessor begins processing user data.
Questions
Questions about subprocessors or data transfers can be sent to [email protected].