Version: 0.12.0
Status: Active
Last Updated: 2026-04-28
- Overview
- Problem Statement
- Target Users
- Design Decisions
- Core Data Model
- Provider Traits
- Authentication & Security
- System Architecture
- Ingestion Layer
- Processing Pipeline
- Retrieval Engine
- API Design
- Database Conventions
- Error Handling
- Testing Strategy
- Observability
- CI/CD Pipeline
- Multi-Tenancy
- Decay & Pruning
- Rate Limiting
- Frontend Architecture
- Audit Log
- Integration Health & DLQ
- MCP Server
- Export & Backup
- Code Quality Standards
- Non-Goals
- Risks & Mitigations
- Milestones
MemoryOps is a Memory Operations Platform designed to give AI agents persistent, structured, and controllable memory. It ingests raw engineering activity from external tools, transforms that activity into typed memory units, and serves optimized context back to agents at query time.
Core abstraction shift: from storage → control.
MemoryOps is not:
- A vector database
- A RAG wrapper
- An agent framework
MemoryOps is:
- The control plane for what AI agents remember
| Problem | Root Cause | Impact |
|---|---|---|
| Session amnesia | No persistence layer | Agents re-ask for context every session |
| Context bloat | Naive prompt stuffing | Token waste, degraded response quality |
| Poor retrieval | Top-K only, no scoring | Irrelevant context pollutes agent output |
| No governance | No lifecycle management | Stale/wrong memories silently affect behavior |
| No debuggability | Black-box retrieval | Can't explain why agent answered incorrectly |
Primary ICP: AI-native engineering teams building:
- Coding agents / dev copilots
- DevOps / SRE agents
- Internal engineering assistants
Qualification criteria:
- Team has an active AI agent in production or development
- Agent uses GitHub and/or Slack as primary tools
- Pain with stateless context is measurable and acknowledged
| # | Decision | Choice | Rationale |
|---|---|---|---|
| 1 | Embedding model | Pluggable EmbeddingProvider trait; fastembed-rs local default |
No external dependency; swap via config |
| 2 | LLM summarization | Pluggable LlmProvider trait; Ollama local default |
Self-hostable by default |
| 3 | Authentication | API key per workspace (X-API-Key header) |
Simple, no OAuth complexity in v0.1 |
| 4 | Retrieval mode | Pull — agents call POST /retrieve |
Clean integration contract |
| 5 | Memory scope | Configurable hierarchy: workspace → agent → user → repo | Flexible without over-engineering |
| 6 | Webhook validation | HMAC-SHA256 via shared WebhookValidator trait |
Consistent across all sources |
Immutable. Written on ingestion, never mutated. Source of truth for all memory lineage.
#[derive(Debug, Clone, Serialize, Deserialize, sqlx::FromRow)]
pub struct RawEvent {
pub id: Uuid,
pub workspace_id: Uuid,
pub source: Source,
pub event_type: EventType,
pub actor: String,
pub payload: serde_json::Value,
pub occurred_at: DateTime<Utc>,
pub ingested_at: DateTime<Utc>,
}
#[derive(Debug, Clone, Serialize, Deserialize, sqlx::Type)]
#[sqlx(type_name = "source", rename_all = "lowercase")]
pub enum Source {
GitHub,
Slack,
Jira,
Linear,
}
#[derive(Debug, Clone, Serialize, Deserialize, sqlx::Type)]
#[sqlx(type_name = "event_type", rename_all = "snake_case")]
pub enum EventType {
PullRequest,
PullRequestReview,
Push,
IssueComment,
Issue,
Message,
Reaction,
}Core product object. Created by the processor. Versioned on edit.
#[derive(Debug, Clone, Serialize, Deserialize, sqlx::FromRow)]
pub struct MemoryUnit {
pub id: Uuid,
pub workspace_id: Uuid,
pub scope: MemoryScope,
pub memory_type: MemoryType,
pub content: String,
pub entities: sqlx::types::Json<Vec<Entity>>,
pub importance_score: f32,
pub importance_overridden: bool, // true if user manually set score
pub source_events: Vec<Uuid>,
pub embedding_id: Option<String>, // Qdrant point ID
pub created_at: DateTime<Utc>,
pub updated_at: DateTime<Utc>,
pub last_accessed_at: Option<DateTime<Utc>>,
pub decay_score: f32,
pub pinned: bool,
pub tags: Vec<String>,
pub version: i32,
pub deleted_at: Option<DateTime<Utc>>, // soft delete
}
#[derive(Debug, Clone, Serialize, Deserialize, sqlx::Type)]
#[sqlx(type_name = "memory_type", rename_all = "lowercase")]
pub enum MemoryType {
Episodic,
Semantic,
}
#### Episodic vs. Semantic Memory
MemoryOps distinguishes between these two core memory types:
* **Episodic Memory**: Represents discrete, point-in-time experiences or events (e.g., a specific GitHub PR merge event, a production deployment failure alert, or a specific Slack observation). They are highly time-sensitive, subject to a mathematical decay rate (fading in relevance over time), and act as raw inputs to the memory engine.
* **Semantic Memory**: Represents durable, consolidated facts, rules, and workspace-level concepts (e.g., "The database connection pool limit is 20", or coding styles, or API usage contracts). Semantic memories do not decay under the normal pipeline, are version-controlled upon modification, and are created either manually via the API or automatically via the Promotion Pipeline (which clusters related episodic memories, summarizes them using an LLM, and promotes the consensus to a semantic memory).#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct MemoryScope {
pub workspace_id: Uuid,
pub agent_id: Option<String>,
pub user_id: Option<String>,
pub repo: Option<String>,
}
impl MemoryScope {
/// Returns a specificity score — used to prefer narrower scopes in retrieval.
pub fn specificity(&self) -> u8 {
let mut score = 0u8;
if self.agent_id.is_some() { score += 4; }
if self.user_id.is_some() { score += 2; }
if self.repo.is_some() { score += 1; }
score
}
}#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct Entity {
pub entity_type: EntityType,
pub value: String,
pub confidence: f32, // 0.0–1.0
}
#[derive(Debug, Clone, Serialize, Deserialize)]
#[serde(rename_all = "snake_case")]
pub enum EntityType {
Person,
Repo,
Branch,
Topic,
File,
Team,
}Every edit to a Semantic memory creates a version row before mutating.
#[derive(Debug, Clone, Serialize, Deserialize, sqlx::FromRow)]
pub struct MemoryVersion {
pub id: Uuid,
pub memory_id: Uuid,
pub workspace_id: Uuid,
pub version: i32,
pub content: String,
pub importance_score: f32,
pub tags: Vec<String>,
pub edited_by: String, // actor (API key name or "system")
pub created_at: DateTime<Utc>,
}Workspace-level configuration is stored in the workspaces.config JSONB column. Missing lifecycle fields are intentionally deserialized as None; the scheduler applies explicit defaults when it runs.
pub struct WorkspaceConfig {
pub promotion_threshold: f32,
pub dedup_cosine_threshold: f32,
pub access_count_trigger: u32,
pub half_life_days: f32,
pub decay_rate_episodic: f32,
pub decay_rate_semantic: f32,
pub llm_provider: Option<String>,
pub embedding_provider: Option<String>,
pub decay_half_life_days: Option<u32>,
pub pruning_threshold: Option<f32>,
}| Field | Default | Validation | Behavior |
|---|---|---|---|
decay_half_life_days |
30 |
Min 1, max 3650 |
When null or absent, the scheduler falls back to the global default of 30 days. |
pruning_threshold |
0.10 |
Min 0.01, max 0.50 |
When null or absent, the scheduler falls back to the global default of 0.10. |
All AI integrations are behind async traits defined in common. No crate outside common imports a concrete provider directly — only the trait.
#[async_trait]
pub trait EmbeddingProvider: Send + Sync + 'static {
async fn embed(&self, text: &str) -> Result<Vec<f32>, ProviderError>;
async fn embed_batch(&self, texts: &[&str]) -> Result<Vec<Vec<f32>>, ProviderError>;
fn dimensions(&self) -> usize;
fn model_name(&self) -> &str;
}Implementations:
| Provider | Crate | Notes |
|---|---|---|
FastEmbedProvider |
fastembed |
Local, no network, default |
OpenAIEmbedProvider |
async-openai |
text-embedding-3-small |
#[async_trait]
pub trait LlmProvider: Send + Sync + 'static {
async fn complete(&self, prompt: &str) -> Result<String, ProviderError>;
async fn summarize(
&self,
text: &str,
max_tokens: usize,
) -> Result<String, ProviderError>;
}Implementations:
| Provider | Notes |
|---|---|
OllamaProvider |
Local HTTP, model configurable, default |
OpenAIProvider |
Chat Completions API |
AnthropicProvider |
Messages API |
#[derive(Debug, thiserror::Error)]
pub enum ProviderError {
#[error("provider request failed: {0}")]
Request(String),
#[error("provider rate limited, retry after {retry_after_secs}s")]
RateLimited { retry_after_secs: u64 },
#[error("provider returned invalid response: {0}")]
InvalidResponse(String),
#[error("provider not configured")]
NotConfigured,
}[embedding]
provider = "fastembed" # fastembed | openai
model = "BAAI/bge-small-en-v1.5"
[llm]
provider = "ollama" # ollama | openai | anthropic
model = "llama3"
base_url = "http://localhost:11434"
timeout_secs = 30
[llm.openai] # only read if provider = "openai"
api_key_env = "OPENAI_API_KEY" # env var name, not value
[llm.anthropic]
api_key_env = "ANTHROPIC_API_KEY"Secret values are never in config files — always resolved from environment variables.
#[derive(Debug, Clone, sqlx::FromRow)]
pub struct ApiKey {
pub id: Uuid,
pub workspace_id: Uuid,
pub name: String,
pub key_hash: String, // Argon2id hash
pub prefix: String, // first 8 chars of plaintext, for display
pub created_at: DateTime<Utc>,
pub last_used_at: Option<DateTime<Utc>>,
pub revoked: bool,
pub revoked_at: Option<DateTime<Utc>>,
}Key format: mops_<workspace_prefix>_<32 random bytes as base58>
Example: YOUR_MEMORYOPS_API_KEY
Hashing: Argon2id with params: m=65536, t=2, p=1 (OWASP recommended)
Key lifecycle:
POST /workspaces/:id/keys→ generate plaintext, hash, store hash, return plaintext once- Client authenticates with
X-API-Key: <plaintext> - On request: hash incoming key, compare to stored hash (constant-time)
- Update
last_used_atasynchronously (non-blocking fire-and-forget) - Revoke: set
revoked = true,revoked_at = now()
Request flows through middleware in this order:
Request
│
▼
TraceLayer # request ID, span creation
│
▼
RequestIdLayer # inject X-Request-ID
│
▼
TimeoutLayer # 30s global timeout
│
▼
CorsLayer # configurable origins
│
▼
AuthLayer # X-API-Key → workspace_id extraction
│ # skipped for /v1/ingest/* and /health
▼
RateLimitLayer # Redis sliding window per workspace
│
▼
Handler
- All DB queries use parameterized statements via
sqlx— no string interpolation workspace_idinjected from authenticated context, never from request body- Webhook payloads validated before deserialization (fail-fast on bad HMAC)
- No secrets in logs —
tracingfields sanitized for key material - Dependency audit via
cargo auditin CI Content-Security-Policy,X-Frame-Options,X-Content-Type-Optionsheaders on all responses- PII in raw payloads never exported (export endpoint strips
payloadfield)
┌───────────────────────────────────────────────────────────────┐
│ MemoryOps │
│ │
│ ┌──────────────────┐ ┌────────────────────────────────┐ │
│ │ Ingestion Svc │ │ Processor Svc │ │
│ │ │ │ │ │
│ │ POST /v1/ingest/ │───▶│ Fast Worker │ Slow Worker │ │
│ │ github | slack │ │ (sync, rules)│ (async, LLM) │ │
│ └──────────────────┘ └────────────────────────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ Postgres Redis Streams │
│ ┌─────────────────┐ ┌──────────────────┐ │
│ │ raw_events │ │ processor_jobs │ │
│ │ memory_units │ │ dlq │ │
│ │ memory_versions│ └──────────────────┘ │
│ │ audit_log │ │ │
│ │ api_keys │◀─────────────────┘ │
│ │ workspaces │ │
│ └─────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────────────────────┐ │
│ │ Retrieval Svc │ │
│ │ Qdrant (semantic) + PG FTS (BM25) │ │
│ │ RRF Fusion → Decay → Promotion │ │
│ └──────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────────────────────┐ │
│ │ API Svc (axum) │◀── Agents / UI │
│ │ Auth → RateLimit → Handlers │ │
│ └──────────────────────────────────────┘ │
└───────────────────────────────────────────────────────────────┘
memoryops/
├── Cargo.toml # workspace root
├── crates/
│ ├── common/ # shared across all crates
│ │ ├── src/
│ │ │ ├── config.rs # AppConfig, loaded from TOML + env
│ │ │ ├── db.rs # PgPool init, migration runner
│ │ │ ├── error.rs # AppError, ProviderError, unified error types
│ │ │ ├── models/ # RawEvent, MemoryUnit, MemoryScope, Entity, ...
│ │ │ ├── providers/
│ │ │ │ ├── mod.rs
│ │ │ │ ├── traits.rs # EmbeddingProvider, LlmProvider traits
│ │ │ │ ├── fastembed.rs
│ │ │ │ ├── ollama.rs
│ │ │ │ ├── openai.rs
│ │ │ │ └── anthropic.rs
│ │ │ └── telemetry.rs # tracing + OTEL init
│ ├── ingestion/
│ │ ├── src/
│ │ │ ├── lib.rs
│ │ │ ├── router.rs # axum routes for /v1/ingest/*
│ │ │ ├── github/
│ │ │ │ ├── mod.rs
│ │ │ │ ├── handler.rs
│ │ │ │ ├── signature.rs
│ │ │ │ └── parser.rs # GitHub payload → RawEvent
│ │ │ └── slack/
│ │ │ ├── mod.rs
│ │ │ ├── handler.rs
│ │ │ ├── parser.rs
│ │ │ └── validator.rs
│ ├── processor/
│ │ ├── src/
│ │ │ ├── lib.rs
│ │ │ ├── worker.rs # Redis consumer loop (fast + slow paths)
│ │ │ ├── extractor.rs # entity extraction
│ │ │ ├── embedder.rs # EmbeddingProvider + Qdrant write
│ │ │ ├── scope.rs # MemoryScope builder
│ │ │ ├── store.rs # DB writes for MemoryUnit
│ │ │ ├── dlq.rs # Dead letter queue
│ │ │ └── pipeline/ # fast/slow path orchestration
│ ├── retrieval/
│ │ ├── src/
│ │ │ ├── lib.rs # router registration
│ │ │ ├── dto.rs # SearchRequest, ListQuery, UpdateMemoryRequest, DTOs
│ │ │ ├── handlers/ # search, list, get, update handlers
│ │ │ ├── search/ # vector, keyword, hybrid (RRF) search modules
│ │ │ ├── promotion/ # decay + eligibility + promotion trigger
│ │ │ ├── access.rs # Redis access counter
│ │ │ └── store.rs # all retrieval DB queries
│ └── api/
│ ├── src/
│ │ ├── main.rs # AppState wiring, router merge, startup
│ │ ├── scheduler.rs # API-owned scheduled background jobs
│ │ ├── middleware/
│ │ │ ├── auth.rs
│ │ │ └── rate_limit.rs
│ │ └── handlers/
│ │ ├── workspaces.rs
│ │ ├── keys.rs
│ │ └── audit.rs
├── frontend/ # React + TypeScript (M5)
│ ├── src/
│ │ ├── components/
│ │ ├── pages/
│ │ ├── hooks/
│ │ ├── api/ # typed API client (generated from OpenAPI)
│ │ └── stores/ # Zustand state
│ └── package.json
├── migrations/ # sqlx numbered migrations
│ ├── 0001_init.sql
│ ├── 0002_ingestion_indexes.sql
│ ├── 0003_processor.sql
│ ├── 0004_retrieval.sql # FTS indexes, access_count column
│ ├── 0005_workspaces.sql
│ ├── 0006_api_keys.sql
│ ├── 0007_audit_log.sql
│ ├── 0008_integrations.sql
│ ├── 0009_retrieval_traces.sql
│ ├── 0010_soft_delete.sql
│ ├── 0011_scheduler.sql
│ ├── 0012_promotion.sql
│ └── 0013_slack.sql
├── docs/
│ ├── SPEC.md
│ ├── FEATURES.md
│ └── openapi.yaml # API contract (source of truth)
├── docker-compose.yml
├── docker-compose.test.yml # isolated test infra
└── .github/
└── workflows/
├── ci.yml
└── release.yml
All shared state is injected via axum State<AppState> — never global statics.
#[derive(Clone)]
pub struct AppState {
pub db: PgPool,
pub redis: ConnectionManager,
pub qdrant: QdrantClient,
pub embedding_provider: Arc<dyn EmbeddingProvider>,
pub llm_provider: Arc<dyn LlmProvider>,
pub config: Arc<AppConfig>,
pub app_secret_key: Arc<Zeroizing<String>>,
}The Ingestion Pipeline is the secure, high-throughput gateway of the MemoryOps control plane. It acts as an HMAC-validated receiver for developer workflow activity, translating raw tool actions into standardized RawEvent structs.
- Security & Signature Validation: Webhooks are never accepted blindly. All incoming requests undergo strict HMAC signature verification using a per-workspace registered integration secret key (e.g.
X-Hub-Signature-256for GitHub,X-Slack-Signaturefor Slack). - Transactional Idempotency: Duplicate webhook deliveries (common in distributed messaging) are rejected automatically before processing. The ingestion transaction guarantees that a raw event is stored in PostgreSQL and enqueued in Redis Streams (
XADD) atomically. - Decoupled Async Architecture: Webhook ingestion is designed for extreme speed, immediately returning a
202 Acceptedstatus to the caller, while asynchronous workers in theprocessorcrate handle parsing, entity extraction, importance scoring, and Qdrant storage.
pub trait WebhookValidator: Send + Sync {
fn validate(
&self,
payload: &[u8],
headers: &HeaderMap,
) -> Result<(), ValidationError>;
}
#[derive(Debug, thiserror::Error)]
pub enum ValidationError {
#[error("missing signature header")]
MissingHeader,
#[error("invalid signature")]
InvalidSignature,
#[error("timestamp too old")] // replay attack prevention
StaleTimestamp,
}GitHubValidator: verifiesX-Hub-Signature-256: sha256=<hmac>SlackValidator: verifiesX-Slack-Signature: v0=<hmac>withv0:{timestamp}:{body}message; rejects timestamps older than 5 minutes
| Event | Actions Captured |
|---|---|
pull_request |
opened, closed, merged, review_requested |
pull_request_review |
submitted (approved, changes_requested, commented) |
push |
all refs; extract commits, author, branch |
issue_comment |
created, edited |
issues |
opened, closed, labeled |
POST /v1/ingest/github/{workspace_id}
│
├─ 1. Validate HMAC → 401 ValidationError on failure
├─ 2. Parse event type from X-GitHub-Event header
├─ 3. Deserialize payload (serde_json) → 400 on parse error
├─ 4. Build RawEvent struct
├─ 5. BEGIN TRANSACTION
│ ├─ INSERT raw_events
│ └─ XADD processor_jobs stream
├─ 6. COMMIT
├─ 7. Return 202 Accepted { event_id }
└─ 8. Async: write audit entry, update integration health
Step 5 uses a Redis pipeline inside the Postgres transaction callback to ensure the job is only enqueued if the DB write succeeds. On transaction rollback, the Redis XADD is not executed.
GitHub may deliver webhooks more than once. Idempotency key = SHA256(source + event_type + payload.id + occurred_at). On conflict, return 202 without re-processing.
| Event | Meaning |
|---|---|
message |
New message in channel or thread |
message.edited |
Message body edited |
reaction_added |
Emoji reaction on a message |
app_mention |
Direct @mention of the MemoryOps app |
Runs in the same process as ingestion (tokio task), processes synchronously before returning the 202.
Steps:
- Extract entities using regex patterns + allowlists
- Assign tags from a rule table (event_type × entity_pattern → tag)
- Score importance from rule table (workspace-configurable weights)
- Write
MemoryUnit(Episodic,embedding_id = None) - Enqueue slow path job
Importance scoring rules (default):
| Event | Default Score |
|---|---|
| Merge to default branch | 0.90 |
| PR review: changes_requested | 0.75 |
| PR opened | 0.60 |
| PR review: approved | 0.55 |
| Issue opened | 0.50 |
| Issue comment | 0.35 |
| Push to feature branch | 0.30 |
| Reaction added | 0.10 |
Long-running tokio task consuming from Redis Streams (XREADGROUP).
XREADGROUP GROUP slow_workers consumer-1 COUNT 10 BLOCK 2000 STREAMS processor_jobs >
│
├─ For each job:
│ ├─ Load MemoryUnit from Postgres
│ ├─ Call LlmProvider::summarize(content)
│ ├─ Call EmbeddingProvider::embed(summary)
│ ├─ Upsert point in Qdrant
│ ├─ Update MemoryUnit: embedding_id, updated_at
│ ├─ Trigger promotion check
│ └─ XACK processor_jobs slow_workers <message_id>
│
└─ On error: increment retry count → DLQ after 3 failures
Consumer group: multiple slow workers can run concurrently. Redis Streams XREADGROUP ensures each job is processed by exactly one worker.
MemoryOps has two live promotion paths.
Access-based promotion (async per retrieval):
1. Load MemoryUnit by id
2. Check Redis access counter (HGET memoryops:access:<id> count)
3. If: memory_type == Episodic
AND importance_score >= promotion_threshold
AND access_count >= access_count_trigger
AND not deleted, not pinned
→ UPDATE memory_type = 'semantic'
4. Log promotion via tracing::info!
Cluster-based promotion (nightly scheduler):
1. Fetch all non-promoted episodic memories per workspace
2. Compute cosine similarity between pairs using Qdrant vectors
3. Group by similarity > dedup_threshold (default 0.92) into clusters
4. For qualifying clusters (avg importance >= promotion_threshold):
a. Select highest-importance unit as canonical
b. LLM summarize concatenated cluster content
c. Write new Semantic MemoryUnit with merged source_events
d. Soft-delete all cluster members
e. Write Qdrant point for new semantic unit
f. Emit AuditEntry: memory_promoted per affected id
POST /v1/workspaces/:id/promote manually triggers the cluster-based pass. The handler uses a Redis SET NX EX 300 distributed lock per workspace to prevent concurrent runs.
Configurable per workspace: promotion_threshold (default: 0.85), access_count_trigger (default: 3), dedup_threshold (default: 0.92).
The retrieval crate supports three search modes on POST /v1/memory/search:
| Mode | Strategy | Notes |
|---|---|---|
vector |
Qdrant ANN search | Requires embedding_id; falls back to empty if provider not configured |
keyword |
PostgreSQL tsvector FTS |
plainto_tsquery('english', query) + GIN index |
hybrid (default) |
RRF fusion of vector + keyword | score = Σ 1/(k + rank) where k=60; normalised 0–1 |
pub const RRF_K: f32 = 60.0;
pub fn rrf_score(rank: u32) -> f32 {
1.0 / (RRF_K + rank as f32)
}Candidates from both legs are fused, scores normalised to [0, 1] by dividing by max raw score, then re-ranked. Vector result is preferred for tie-breaking.
pub fn decay_score(importance_score: f32, elapsed_secs: f64, half_life_secs: f64) -> f32 {
// decay_score = importance_score × 0.5^(elapsed / half_life)
}Default half-life: 30 days. Applied in bulk via apply_decay_scores_with_half_life with the workspace-resolved half-life value; skips pinned and importance-overridden units.
Every search result triggers:
access::record_access(redis, memory_id)— increments Redis HINCR counter (TTL 90 days)store::touch_last_accessed(db, id)— async update oflast_accessed_atpromotion::check_and_promote(state, workspace_id, result_ids)— async promotion eligibility check
GET /v1/memory supports full filtering: workspace_id, memory_type, pinned, min_importance, sort by importance_score | decay_score | updated_at | created_at, direction asc | desc, and cursor pagination.
PATCH /v1/memory/:id supports: content, pinned, importance_score (sets importance_overridden = true), tags. Validates importance_score ∈ [0.0, 1.0].
Full per-component score breakdown (semantic_similarity, importance, recency, source_authority, memory_type_weight) is tracked in RetrievalTraceEntry. POST /v1/retrieve is live as of M6 and returns packed memories with score breakdowns and trace data. Retrieval traces are persisted for 30 days in the retrieval_traces table and can be inspected through GET /v1/retrieve/trace/:query_id.
- Versioning: All routes prefixed
/v1/ - Auth:
X-API-Keyheader required on all routes except/v1/ingest/*,/health, andPOST /v1/workspaces(which requiresx-admin-token) - Content-Type:
application/jsonfor all request/response bodies - Pagination: cursor-based on list endpoints (
?after=<cursor>&limit=<n>, default limit 20, max 100) - Errors: unified error envelope (see §14)
- Idempotency:
POSTendpoints accept optionalIdempotency-Keyheader - Timestamps: all
DateTimefields in ISO 8601 UTC (2026-04-27T18:00:00Z) - IDs: all IDs are UUIDs v7 (time-ordered, sortable)
{
"error": {
"code": "memory_not_found",
"message": "Memory unit with id 'abc...' not found"
}
}| Method | Path | Status | Description |
|---|---|---|---|
| POST | /v1/ingest/github/{workspace_id} |
✅ Live | GitHub webhook receiver |
| POST | /v1/ingest/slack/{workspace_id} |
✅ Live | Slack Events API receiver |
| POST | /v1/ingest/linear/{workspace_id} |
✅ Live | Linear webhook receiver |
| POST | /v1/ingest/jira/{workspace_id} |
✅ Live | Jira Cloud admin webhook receiver |
| Method | Path | Status | Description |
|---|---|---|---|
| GET | /v1/memory |
✅ Live | List (filter by type/pinned/importance/tag; sort; paginated) |
| GET | /v1/memory/:id |
✅ Live | Get with full entity, scope, score |
| PATCH | /v1/memory/:id |
✅ Live | Update: content, pin, tag, importance override |
| DELETE | /v1/memory/:id |
✅ Live | Soft delete |
| POST | /v1/memory/search |
✅ Live | Hybrid/vector/keyword search with RRF fusion |
| POST | /v1/memory/:id/promote |
✅ Live | Force episodic → semantic |
| POST | /v1/memory/:id/restore |
✅ Live | Restore soft-deleted memory |
| POST | /v1/memory/bulk |
✅ Live | Bulk pin / bulk delete |
| GET | /v1/memory/:id/history |
✅ Live | Version history |
| POST | /v1/memory/merge |
✅ Live | Merge two semantic memory units |
| Method | Path | Status | Description |
|---|---|---|---|
| POST | /v1/retrieve |
✅ Live | Core retrieval with token packing + trace |
| GET | /v1/retrieve/trace/:query_id |
✅ Live | Retrieval trace |
| Method | Path | Status | Description |
|---|---|---|---|
| POST | /v1/workspaces |
✅ Live | Create workspace |
| GET | /v1/workspaces/:id |
✅ Live | Get workspace |
| PATCH | /v1/workspaces/:id/config |
✅ Live | Update config |
| POST | /v1/workspaces/:id/promote |
✅ Live | Manual promotion pass (workspace-scoped lock) |
| POST | /v1/workspaces/:id/integrations |
✅ Live | Add integration |
| GET | /v1/workspaces/:id/integrations |
✅ Live | List with health |
| DELETE | /v1/workspaces/:id/integrations/:source |
✅ Live | Remove integration |
| Method | Path | Status | Description |
|---|---|---|---|
| POST | /v1/workspaces/:id/keys |
✅ Live | Create key (plaintext returned once) |
| GET | /v1/workspaces/:id/keys |
✅ Live | List keys |
| DELETE | /v1/workspaces/:id/keys/:key_id |
✅ Live | Revoke key |
| Method | Path | Status | Description |
|---|---|---|---|
| GET | /v1/workspaces/:id/audit |
✅ Live | Paginated audit log |
| GET | /v1/workspaces/:id/dlq |
✅ Live | List DLQ jobs |
| POST | /v1/workspaces/:id/dlq/:job_id/retry |
✅ Live | Retry DLQ job |
| DELETE | /v1/workspaces/:id/dlq/:job_id |
✅ Live | Discard DLQ job |
| Method | Path | Status | Description |
|---|---|---|---|
| GET | /v1/workspaces/:id/export |
✅ Live | Stream JSONL export |
| Method | Path | Description |
|---|---|---|
| GET | /health |
Liveness (no auth) |
| GET | /health/ready |
Readiness: checks DB + Redis + Qdrant |
- All migrations in
migrations/numbered sequentially:0001_init.sql - Migrations are additive only — never drop columns, never rename in-place
- Deprecate columns by adding
_deprecatedsuffix and a TODO migration - All tables include
created_at TIMESTAMPTZ NOT NULL DEFAULT now() - All tables include
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()+ trigger - Soft deletes use
deleted_at TIMESTAMPTZ(NULL = not deleted) - All foreign keys have explicit
ON DELETEbehavior defined
-- raw_events
CREATE INDEX idx_raw_events_workspace_occurred ON raw_events(workspace_id, occurred_at DESC);
CREATE UNIQUE INDEX idx_raw_events_idempotency ON raw_events(idempotency_key);
-- memory_units
CREATE INDEX idx_memory_units_workspace_type ON memory_units(workspace_id, memory_type);
CREATE INDEX idx_memory_units_decay ON memory_units(workspace_id, decay_score) WHERE deleted_at IS NULL;
CREATE INDEX idx_memory_units_scope ON memory_units(workspace_id, agent_id, user_id, repo);
CREATE INDEX idx_memory_units_tags ON memory_units USING gin(tags);
-- M4 additions:
CREATE INDEX idx_memory_units_fts ON memory_units USING GIN(to_tsvector('english', content));
CREATE INDEX idx_memory_units_workspace_type_score ON memory_units(workspace_id, memory_type, importance_score DESC) WHERE deleted_at IS NULL;
CREATE INDEX idx_memory_units_workspace_pinned ON memory_units(workspace_id, pinned, updated_at DESC) WHERE deleted_at IS NULL AND pinned = true;
-- audit_log
CREATE INDEX idx_audit_log_workspace_time ON audit_log(workspace_id, occurred_at DESC);CREATE OR REPLACE FUNCTION set_updated_at()
RETURNS TRIGGER AS $$
BEGIN
NEW.updated_at = now();
RETURN NEW;
END;
$$ LANGUAGE plpgsql;
CREATE TRIGGER trg_memory_units_updated_at
BEFORE UPDATE ON memory_units
FOR EACH ROW EXECUTE FUNCTION set_updated_at();- All queries receive
workspace_idas a bind param — never interpolated - Use
sqlx::query_as!macro for compile-time SQL checking - Use
RETURNING *on inserts to avoid round-trips - Never
SELECT *in application code — always explicit column lists - Connection pool:
max_connections = 20(configurable),min_connections = 2
#[derive(Debug, thiserror::Error)]
pub enum AppError {
#[error("not found: {resource}")]
NotFound { resource: String },
#[error("unauthorized")]
Unauthorized,
#[error("forbidden: insufficient scope")]
Forbidden,
#[error("validation error: {0}")]
Validation(String),
#[error("conflict: {0}")]
Conflict(String),
#[error("rate limited")]
RateLimited { retry_after_secs: u64 },
#[error("provider error: {0}")]
Provider(#[from] ProviderError),
#[error("database error: {0}")]
Database(#[from] sqlx::Error),
#[error("internal error")]
Internal(#[from] anyhow::Error),
}impl IntoResponse for AppError {
fn into_response(self) -> Response {
let (status, code) = match &self {
AppError::NotFound { .. } => (StatusCode::NOT_FOUND, "not_found"),
AppError::Unauthorized => (StatusCode::UNAUTHORIZED, "unauthorized"),
AppError::Forbidden => (StatusCode::FORBIDDEN, "forbidden"),
AppError::Validation(_) => (StatusCode::BAD_REQUEST, "validation_error"),
AppError::Conflict(_) => (StatusCode::CONFLICT, "conflict"),
AppError::RateLimited{..} => (StatusCode::TOO_MANY_REQUESTS, "rate_limited"),
AppError::Provider(_) => (StatusCode::BAD_GATEWAY, "provider_error"),
AppError::Database(_)
| AppError::Internal(_) => (StatusCode::INTERNAL_SERVER_ERROR, "internal_error"),
};
if status == StatusCode::INTERNAL_SERVER_ERROR {
tracing::error!(error = ?self, "internal error");
}
Json(json!({
"error": {
"code": code,
"message": self.to_string(),
}
})).into_response()
}
}- Never
unwrap()orexpect()in non-test code — use?propagation - Never expose internal error details (DB messages, stack traces) to clients
- All errors logged with
tracing::error!at the handler boundary, not deep in call stack anyhow::Contextused to add context when converting errors across crate boundaries- Panics are treated as bugs —
tokio::spawntasks catch panics and emit error spans
┌──────────────┐
│ E2E Tests │ (few, slow) docker-compose.test.yml
├──────────────┤
│ Integration │ (moderate) real DB + Redis + Qdrant
├──────────────┤
│ Unit Tests │ (many, fast) pure logic, no I/O
└──────────────┘
- Location:
#[cfg(test)]module at bottom of each source file - Coverage targets:
- RRF fusion: rank ordering, score normalisation, tie-breaking
- Decay formula: half-life boundary conditions
- Keyword search builder: filter pushdown SQL correctness
- Promotion eligibility: all flag combinations
- Webhook validator: valid HMAC, invalid HMAC, stale timestamp
- MemoryScope specificity ordering
- Location:
tests/directory per crate - Use
sqlx::testmacro for DB tests — each test gets a fresh schema - Use
wiremockfor mocking LLM and embedding provider HTTP calls - Live-service tests tagged
#[ignore]— requiredocker-compose.test.yml
- Spin up full stack via
docker-compose.test.yml - Cover: ingest → process → search → verify memory appears in results
- Cover: promotion trigger via repeated access
- Run on every PR, block merge on failure
- Use
proptestfor token packing algorithm — verify invariants:- Packed tokens never exceed
token_budget - No two included memories have similarity >
dedup_threshold - All excluded memories appear in trace
- Packed tokens never exceed
- Minimum 80% line coverage enforced in CI via
cargo-llvm-cov - Coverage report uploaded as CI artifact
- All logs via
tracingcrate — noprintln!oreprintln!in application code - Log format: JSON in production, pretty-printed in development
- Required fields on every request span:
request_id,workspace_id,method,path,status_code,latency_ms
| Metric | Type | Labels |
|---|---|---|
memoryops_ingest_events_total |
Counter | source, event_type, status |
memoryops_processor_job_duration_ms |
Histogram | path (fast/slow) |
memoryops_retrieval_duration_ms |
Histogram | mode (vector/keyword/hybrid) |
memoryops_memory_units_total |
Gauge | workspace_id, memory_type |
memoryops_decay_updated_total |
Counter | workspace_id |
memoryops_dlq_jobs_total |
Gauge | status (pending/failed) |
memoryops_provider_latency_ms |
Histogram | provider, operation |
// GET /health/ready
// Returns 200 only if all dependencies are healthy
pub async fn readiness() -> impl IntoResponse {
let (database, redis, qdrant) = tokio::join!(
check_database(),
check_redis(),
check_qdrant()
);
let ready = database.is_ready() && redis.is_ready() && qdrant.is_ready();
let status = if ready { StatusCode::OK } else { StatusCode::SERVICE_UNAVAILABLE };
(status, Json(json!({ "status": if ready { "ok" } else { "unavailable" }, "checks": { ... } })))
}Triggered on: every push, every PR.
jobs:
check:
- cargo fmt --check
- cargo clippy -- -D warnings
- cargo audit
test:
- cargo test --workspace
- cargo llvm-cov --lcov --output-path lcov.info
- Upload coverage artifact
integration:
services: [postgres, redis, qdrant]
- cargo test --workspace -- --ignoredmain— always deployable; protected, requires PR + CI passfeature/*— feature branches; squash-merged to mainfix/*— bug fixes- No long-lived branches
Follows Conventional Commits:
feat(retrieval): add hybrid RRF search
fix(processor): correct decay half-life formula
chore(deps): update axum to 0.8
docs(spec): update milestones — UI moved to M5
- Every table includes
workspace_id UUID NOT NULL workspace_idis always a bind param — enforced atsqlxlevel, not application layer- No cross-workspace queries possible — no global list endpoints without workspace scope
MemoryScopeprovides sub-workspace granularity (agent / user / repo)- Row-level isolation verified by integration tests that assert workspace A cannot access workspace B data
decay_score(t) = importance_score × 0.5^(elapsed_secs / half_life_secs)
| Parameter | Default | Configurable |
|---|---|---|
decay_half_life_days |
30 | Per workspace via WorkspaceConfig |
pruning_threshold |
0.10 | Per workspace via WorkspaceConfig |
| Pinned memories | Skipped entirely | N/A |
importance_overridden |
Skipped | N/A |
Scheduler behavior:
- The scheduler iterates active workspaces and reads each workspace's
WorkspaceConfigfrom theworkspaces.configJSONB column. decay_half_life_dayscontrols the half-life passed into the decay formula for that workspace.pruning_thresholdcontrols the soft-delete cutoff for that workspace.- If a workspace config is null, missing either lifecycle field, or cannot be parsed, the scheduler logs a warning and falls back to
30days and0.10. - Processes all non-pinned, non-overridden units in workspace-scoped UPDATE batches.
- Runs daily at 02:00 UTC (scheduler — M6+)
- Hard delete: separate job runs 30 days after
deleted_at
Implemented via Redis sliding window counter.
Key pattern: rate:{workspace_id}:{endpoint_group}:{window_start_unix}
Algorithm: INCR + EXPIREAT per window
All limits configurable per workspace. Exceeding returns 429 Too Many Requests.
| Group | Routes | Limit |
|---|---|---|
ingest |
/v1/ingest/* |
300 RPM per workspace |
memory |
/v1/memory/* |
120 RPM per workspace |
api |
All other /v1/* |
120 RPM per workspace |
Window: 60s sliding. Excess → 429 Too Many Requests with Retry-After.
| Concern | Choice | Reason |
|---|---|---|
| Framework | React 19 + TypeScript | Stable, wide ecosystem |
| Build | Vite | Fast HMR, ESM-native |
| State | Zustand | Minimal, no boilerplate |
| Data fetching | TanStack Query | Cache + stale-while-revalidate |
| Styling | Tailwind CSS v4 | Utility-first, consistent |
| Components | shadcn/ui | Accessible, composable |
| Tables | TanStack Table | Powerful, headless |
| Charts | Recharts | Lightweight, declarative |
| Testing | Vitest + React Testing Library | Fast, co-located |
| E2E | Playwright | Cross-browser |
| Route | Component | Live Endpoints Used |
|---|---|---|
/ |
Dashboard | GET /health/ready, GET /v1/memory (counts) |
/memory |
MemoryExplorer | GET /v1/memory, POST /v1/memory/search, PATCH /v1/memory/:id |
/memory/:id |
MemoryDetail | GET /v1/memory/:id, PATCH /v1/memory/:id |
/ingest |
WebhookTester | POST /v1/ingest/github/{workspace_id} (dev tool) |
/settings |
WorkspaceSettings | config display only (M6 write) |
Views for /retrieve/trace, /lifecycle, /audit, /integrations are stubbed with empty states in M5 and wired to real endpoints in M6+.
// Global store — only for cross-cutting concerns
interface AppStore {
workspaceId: string;
apiKey: string; // stored in memory only, never localStorage
setWorkspace: (id: string, key: string) => void;
}
// Server state — TanStack Query handles caching
const { data: memories } = useQuery({
queryKey: ['memories', workspaceId, filters],
queryFn: () => MemoryApi.list({ workspaceId, ...filters }),
staleTime: 30_000,
});Security: API key stored in memory (Zustand store) only. Never written to localStorage, sessionStorage, or cookies. Cleared on tab close.
- No business logic in components — logic lives in custom hooks
- All data fetching in hooks, not components
- Every interactive component has a
data-testidattribute for testing - No inline styles — Tailwind only
- Accessibility: all interactive elements have ARIA labels; keyboard-navigable
#[derive(Debug, Clone, sqlx::FromRow)]
pub struct AuditEntry {
pub id: Uuid,
pub workspace_id: Uuid,
pub actor: String, // API key name or "system" or "scheduler"
pub action: AuditAction,
pub target_id: Uuid,
pub target_type: String,
pub diff: Option<serde_json::Value>,
pub occurred_at: DateTime<Utc>,
}Audit writes are fire-and-forget (tokio::spawn) — never block the primary request path.
pub struct IntegrationHealth {
pub workspace_id: Uuid,
pub source: Source,
pub last_event_at: Option<DateTime<Utc>>,
pub events_24h: i64,
pub errors_24h: i64,
pub status: IntegrationStatus,
}- Backend: Redis list
dlq:{workspace_id} - Auto-retry: 3 attempts with exponential backoff (1s, 4s, 16s)
- After 3 failures: written to DLQ
- DLQ entries expire after 7 days
M11 introduces a dedicated crates/mcp/ server crate that exposes MemoryOps retrieval and storage workflows through the Model Context Protocol 2025-06-18 specification. MCP is a separate transport from the REST API and is not represented in docs/openapi.yaml.
| Concern | Decision |
|---|---|
| Default port | 3003 |
| Port env var | MCP_PORT |
| Transport env var | MCP_TRANSPORT=stdio|http|sse, default stdio; compose sets http for local clients |
| Dependency boundary | mcp depends on common, retrieval, and processor; it must not depend on the api crate |
| Transport | Use Case | Contract |
|---|---|---|
| stdio | Local agents and editor-launched processes | Newline-delimited JSON-RPC 2.0 over stdin/stdout; stdout is flushed after every response |
| HTTP/SSE | Local, private-network, or tunnelled MCP clients | HTTP Streamable uses /mcp; legacy SSE uses /mcp/sse. Do not expose MCP publicly without explicit auth and network controls. |
MCP clients authenticate during the initialize request by passing a workspace API key as a Bearer token in _meta.auth.token:
{
"jsonrpc": "2.0",
"id": 1,
"method": "initialize",
"params": {
"protocolVersion": "2025-06-18",
"_meta": {
"auth": {
"token": "Bearer YOUR_MEMORYOPS_API_KEY"
}
}
}
}The MCP crate validates the key against the same api_keys table and Argon2id verification logic used by REST middleware. On success, the resolved workspace_id is stored in session context and injected into subsequent tool calls. Missing, invalid, or revoked keys return JSON-RPC error code -32001 with message Unauthorized.
For HTTP SSE transport, POST /mcp also accepts Authorization: Bearer <api-key> so browser-based and hosted MCP clients can authenticate each request without relying on a long-lived stdio session.
The MCP server exposes exactly three tools. Tool inputs never accept workspace_id; the workspace is always taken from the authenticated MCP session.
Retrieves token-packed memory context using in-process hybrid retrieval.
Input schema:
{
"type": "object",
"required": ["query"],
"properties": {
"query": { "type": "string" },
"limit": { "type": "integer", "minimum": 1, "maximum": 50, "default": 10 },
"min_score": { "type": "number", "minimum": 0.0, "default": 0.0 }
},
"description": "workspace_id is injected from the authenticated MCP session."
}Output schema:
{
"type": "array",
"items": {
"type": "object",
"required": ["id", "content", "memory_type", "tags", "score", "importance_score", "created_at", "source"],
"properties": {
"id": { "type": "string", "format": "uuid" },
"content": { "type": "string" },
"memory_type": { "type": "string", "enum": ["episodic", "semantic"] },
"tags": { "type": "array", "items": { "type": "string" } },
"score": { "type": "number" },
"importance_score": { "type": "number" },
"created_at": { "type": "string", "format": "date-time" },
"source": { "type": "string" }
}
}
}Searches memory units without token packing.
Input schema:
{
"type": "object",
"required": ["query"],
"properties": {
"query": { "type": "string" },
"search_type": { "type": "string", "enum": ["hybrid", "keyword", "vector"], "default": "hybrid" },
"limit": { "type": "integer", "minimum": 1, "maximum": 100, "default": 20 },
"filters": {
"type": "object",
"properties": {
"tags": { "type": "array", "items": { "type": "string" } },
"memory_type": { "type": "string", "enum": ["episodic", "semantic"] }
}
}
}
}Output schema: same memory unit array shape as memory_retrieve.
Stores an agent-authored episodic memory and enqueues slow-path processing through the same Redis stream used by ingestion.
Input schema:
{
"type": "object",
"required": ["content"],
"properties": {
"content": { "type": "string" },
"source": { "type": "string", "default": "mcp" },
"tags": { "type": "array", "items": { "type": "string" } },
"importance": { "type": "number", "minimum": 0.0, "maximum": 1.0, "default": 0.5 }
}
}Output schema:
{
"type": "object",
"required": ["id", "created_at"],
"properties": {
"id": { "type": "string", "format": "uuid" },
"created_at": { "type": "string", "format": "date-time" }
}
}GET /v1/workspaces/:id/export— streams JSONL, one memory unit per line- Uses chunked transfer encoding — no in-memory buffer for large workspaces
- Excludes:
payloadfrom raw events (may contain PII), embeddings (raw vectors)
- Format:
rustfmtenforced in CI - Lint:
clippywith-D warnings; all warnings treated as errors in CI - No
unwrap()/expect()outside tests - No
clone()on large structs in hot paths — preferArc<T> - Dead code:
#[allow(dead_code)]is forbidden - Unsafe: forbidden unless in a dedicated file with safety comment
- MSRV: Rust 1.88.0 stable, pinned in
rust-toolchain.toml
- Strict mode:
"strict": trueintsconfig.json - Format: Prettier enforced in CI
- Lint: ESLint with
@typescript-eslint+react-hooksrules - No
any: Useunknown+ type guards
- All queries use parameterized binding — no string interpolation ever
- Explain-analyze run on all new queries before merge
- Migrations tested in integration suite before merge
- Generic "second brain" or consumer product
- Full agent runtime or framework
- Vector database replacement
- Real-time streaming ingestion
- Custom embedding model fine-tuning
- Push-mode context injection (agent SDK — v0.3+)
- SSO / OAuth login (v0.3+ SaaS phase)
- Multi-region deployment
- Import from export (v0.3)
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| LLM latency blocks throughput | Medium | High | Async-only slow path; local Ollama default |
| Large context windows reduce retrieval need | Low | Medium | Moat is lifecycle + governance, not retrieval size |
| Retrieval commoditized by frameworks | Medium | Medium | Differentiator is ingestion + control UI + trace |
| GitHub/Slack API changes break parsers | Low | High | Versioned event types; parser tests against fixtures |
| Qdrant unavailability degrades retrieval | Low | High | Fallback to keyword-only if Qdrant unreachable |
| Webhook replay attacks | Low | Medium | Idempotency key + timestamp validation (5min window) |
| Postgres migration failure in prod | Low | Critical | Always run migrations in staging first; rollback plan per migration |
| # | Deliverable | Key Acceptance Criteria | Status |
|---|---|---|---|
| M1 | Rust workspace + docker-compose + migrations scaffold | cargo build clean; docker compose up starts all infra |
✅ Complete |
| M2 | GitHub webhook ingestion + RawEvent writes | Webhook delivers → raw_events row exists; idempotency works |
✅ Complete |
| M3 | Fast path processor + episodic MemoryUnit | Ingest PR → MemoryUnit with entities + importance created | ✅ Complete |
| M4 | Retrieval crate — search, list, get, update, decay, promotion | POST /v1/memory/search returns ranked results; hybrid RRF works; decay pass runs; cargo test passes |
✅ Complete |
| M5 | React Memory Control Center (against live API) | Explorer, detail, search views functional; webhook tester fires real ingestion | ✅ Complete |
| M6 | Full REST API + auth + rate limiting + audit + soft delete | All endpoints passing integration tests; 401/429 enforced; POST /v1/retrieve live |
✅ Complete |
| M7 | Slow path worker + embeddings + Qdrant write | MemoryUnit gets embedding_id; Qdrant point queryable; vector leg of hybrid search active |
✅ Complete |
| M8 | Promotion pipeline (batch clustering) | Episodic cluster → Semantic MemoryUnit after threshold | ✅ Complete |
| M9 | Slack ingestion | Slack message → MemoryUnit via same pipeline | ✅ Complete |
| M10 | Linear + Jira ingestion | Linear/Jira webhooks validate signatures, normalize supported events to RawEvent, enqueue processor jobs, and produce MemoryUnits with source-specific scoring/entities | ✅ Complete |
| M11 | MCP server | crates/mcp/ exposes memory_retrieve, memory_search, and memory_store over stdio and HTTP SSE with workspace API key auth |
✅ Complete |
| M12 | Lifecycle configuration | Workspace config controls decay half-life and pruning threshold per workspace | ✅ Complete |