This guide details the deployment strategies, configuration options, and architectural practices for running MemoryOps in production environments, such as Docker, Kubernetes (K8s), or K3s.
Pre-built multi-platform images (linux/amd64, linux/arm64) are published to Docker Hub:
| Image | Tag | Description |
|---|---|---|
quazmoz/memoryops |
api-latest |
API server (Rust/axum) |
quazmoz/memoryops |
mcp-latest |
MCP gateway (Rust) |
quazmoz/memoryops |
frontend-latest |
Control UI (React/nginx) |
Versioned tags follow the pattern api-0.1.0, mcp-0.1.0, frontend-0.1.0.
# Pull all images
docker pull quazmoz/memoryops:api-latest
docker pull quazmoz/memoryops:mcp-latest
docker pull quazmoz/memoryops:frontend-latestIn a production environment, the MemoryOps services are decoupled into stateless application layers and stateful persistence layers to ensure high availability, horizontal scalability, and resilience.
┌───────────────────────┐
│ Ingress / Load │
│ Balancer │
└──────────┬────────────┘
│ HTTP / gRPC
┌───────────────────┼───────────────────┐
▼ ▼ ▼
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ API Pod 1 │ │ API Pod 2 │ │ API Pod N │ (Stateless API replicas)
└────────┬────────┘ └────────┬────────┘ └────────┬────────┘
│ │ │
└─────────────┬─────┴─────────────┬─────┘
│ Postgres │ Redis / Qdrant
▼ ▼
┌───────────────────┐ ┌───────────────────┐
│ Postgres Database │ │ Redis / Qdrant │ (Stateful Storage / Vector Search)
│ Cluster │ │ Services │
└───────────────────┘ └───────────────────┘
- Stateless API Layer (
api): The core Rust application containing the Axum web server and worker pools. These processes handle ingest, query, and background processors (clustering, decay). They do not maintain state on local disk and can scale horizontally. - Stateless MCP Gateway (
mcp): Optional gateway facilitating communication using Model Context Protocol (MCP) transport. - Stateless Frontend Layer (
frontend): The static SPA bundled with Vite, served via an embedded Nginx/web server container. - Stateful Services Layer: Postgres (relational/metadata storage), Redis (event queue and job locking), and Qdrant (vector search database). These should be backed by persistent storage or managed cloud services (e.g., AWS RDS, ElastiCache, Qdrant Cloud).
When deploying to Kubernetes or K3s, follow these best practices for replica scaling and database operations.
By default, the backend API container attempts to run database migrations on boot. In a horizontally scaled deployment, starting multiple replicas simultaneously causes migrations to race and fail due to lock contention.
To prevent this:
- Set the environment variable
SKIP_MIGRATIONS=trueon your main API Deployment specs. - Run database migrations exactly once per deployment using a Kubernetes Job or an Init Container that executes before the API pods launch.
Example Migration Job Spec (memoryops-db-migrate.yaml):
apiVersion: batch/v1
kind: Job
metadata:
name: memoryops-db-migrate
namespace: memoryops
spec:
template:
spec:
restartPolicy: OnFailure
containers:
- name: migrate
image: quazmoz/memoryops:api-latest
command: ["/usr/local/bin/api"] # Triggers migrations when SKIP_MIGRATIONS is false
env:
- name: DATABASE_URL
valueFrom:
secretKeyRef:
name: memoryops-secrets
key: database-url
- name: SKIP_MIGRATIONS
value: "false" # Explicitly run migrations in this job
- name: REDIS_URL
value: "redis://memoryops-redis:6379"
- name: QDRANT_URL
value: "http://memoryops-qdrant:6334"
- name: APP_SECRET_KEY
valueFrom:
secretKeyRef:
name: memoryops-secrets
key: app-secret-key
- name: WORKSPACE_CREATION_SECRET
valueFrom:
secretKeyRef:
name: memoryops-secrets
key: workspace-creation-secretSince the Agent Library is database-backed (stored in the versioned agent_resources tables, with agent_skills retained for compatibility), individual API replicas do not require local persistent volumes (PVs) or directory mounts for .gemini/skills or .claude/skills.
You can safely set the replicas count to 2+ in your Deployment manifest.
Example API Deployment Spec (memoryops-api-deployment.yaml):
apiVersion: apps/v1
kind: Deployment
metadata:
name: memoryops-api
namespace: memoryops
spec:
replicas: 3
selector:
matchLabels:
app: memoryops-api
template:
metadata:
labels:
app: memoryops-api
spec:
containers:
- name: api
image: quazmoz/memoryops:api-latest
ports:
- containerPort: 8080
readinessProbe:
httpGet:
path: /health/ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 15
periodSeconds: 20
env:
- name: SKIP_MIGRATIONS
value: "true" # Skip migrations on boot for replicas
- name: DATABASE_URL
valueFrom:
secretKeyRef:
name: memoryops-secrets
key: database-url
- name: REDIS_URL
value: "redis://memoryops-redis:6379"
- name: QDRANT_URL
value: "http://memoryops-qdrant:6334"
- name: QDRANT_CHECK_COMPATIBILITY
value: "false" # Bypass client-server compatibility check
- name: APP_SECRET_KEY
valueFrom:
secretKeyRef:
name: memoryops-secrets
key: app-secret-key
- name: WORKSPACE_CREATION_SECRET
valueFrom:
secretKeyRef:
name: memoryops-secrets
key: workspace-creation-secret
- name: APP_ENV
value: "production"
- name: WORKSPACE_CREATION_ENABLED
value: "false"
- name: TRUSTED_PROXY_CIDRS
value: "10.0.10.0/24" # Use only ingress/reverse-proxy CIDRs, not the whole VPCWhen the environment variable APP_ENV is set to production, the backend enforces strict validation on startup:
- Secret Key Validation: The API will crash on boot if
APP_SECRET_KEYis set to the development fallback valuedev-placeholder. If workspace creation is enabled,WORKSPACE_CREATION_SECRETmust also be a real value. - Workspace Creation Switch: Set
WORKSPACE_CREATION_ENABLED=falseafter initial bootstrap soPOST /v1/workspacesis rejected even if the admin token leaks. - Database & Encryption: These values must be long, randomly generated secrets securely stored in a KMS (Key Management Service) or Kubernetes Secret and injected at runtime.
By default, MemoryOps rejects custom skill tools targeting private IP addresses (such as RFC-1918 blocks, link-local metadata endpoints, and loopback) to prevent Server-Side Request Forgery (SSRF).
If you deploy MemoryOps in a private VPC or behind a secure tunnel (e.g. Cloudflare Tunnels, tailscale) and need the server to call tool endpoints running on internal IPs:
- In
config.tomlunder[server], configure:allow_private_ips = true
- Or set the environment variable:
MEMORYOPS_ALLOW_PRIVATE_IPS=true
Warning
Only enable allow_private_ips in secure, single-tenant, private networking environments. In public multi-tenant deployments, keep it set to false.
When behind an Ingress Controller (like NGINX, Traefik, or an AWS ALB), the user's IP is forwarded using the X-Forwarded-For (XFF) header. To prevent IP spoofing, configure TRUSTED_PROXY_CIDRS with the CIDR ranges of your ingress controllers. Only headers sent from these CIDRs will be parsed.
For running in plain Docker environments, MemoryOps uses a multi-file composition approach.
- Development/Local (
docker-compose.yml): Binds ports for databases (5432, 6379, 6334) to loopback127.0.0.1so you can connect local developer tools directly. - Production Hardened Overlay (
docker-compose.prod.yml): Modifies the base compose layout for production.
Run the composition by merging the configurations:
docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d- Network Isolation: The databases (postgres, redis, qdrant) have their port bindings completely removed using the
!reset []directive. This prevents them from binding to host interfaces, isolating them strictly within the Docker bridge network. - Host Binding Limits: The
api,mcp, andfrontendcontainers bind their listening ports strictly to loopback127.0.0.1. - Reverse Proxy Dependency: Place a reverse proxy (e.g., Caddy, NGINX, or Traefik) directly on the host machine to handle SSL termination, routing traffic from the public network to the internal loopback ports (
8080for API,5173for Frontend).
In a remote server setup, you must configure how the API server reaches your LLM and Embedding models. See LLM & Embedding Provider Configuration for individual provider parameters.
If you host Ollama in a separate container or on a GPU-enabled host machine:
- Docker Desktop: Set
llm.base_url = "http://host.docker.internal:11434" - Linux Container Engine: Since
host.docker.internalis not automatically mapped on Linux, add the host gateway resolution in your compose configuration:services: api: extra_hosts: - "host.docker.internal:host-gateway"
- Remote VM / Shared Cluster: Set
llm.base_urlto the private domain name or IP of your GPU instance (e.g.,http://ollama-service.internal:11434).
If you choose to use the local fastembed provider (provider = "fastembed"), the Rust API downloads the BAAI BGE model on first launch. If your production pods run in a restricted or offline network:
- Consider switching to
openaiembeddings (provider = "openai") which query a remote HTTPS API. - Or prepopulate the model cache directory inside your Docker image during build time (defaulting to the cache dir of
fastembed-rs).
Agent skills, agent profiles, prompts, and reusable instructions are stored directly in PostgreSQL (agent_resources and agent_resource_versions) scoped by workspace_id. This guarantees database consistency, preserves immutable version history, and removes file synchronization issues across stateless API replicas. The legacy agent_skills API remains available for Claude/Gemini skill sync workflows.
When a workspace is created, or when listing/retrieving an Agent Library kind that has 0 resources, the server seeds safe starter resources without overwriting existing rows. Skill defaults come from the server filesystem's .gemini/skills/ and .claude/skills/ directories; prompts, agent profiles, and reusable instructions are seeded from built-in MemoryOps defaults with version history.
The default skill markdown files are packed into the production Docker image during the build stage (COPY .gemini /app/.gemini and COPY .claude /app/.claude).
Because resources are in the Postgres database, modifying files inside your local workspace's .gemini/skills or .claude/skills directory will not automatically update a remote server. You can synchronize skill changes bidirectionally:
From your workstation or a deploy script, run the Node.js helper to sync local skills to/from the remote server:
# Sync local markdown files to the remote Postgres database
API_KEY=<your-workspace-api-key> node scripts/memoryops-client.js sync-skillsThe VS Code extension includes the Sync Agent Skills command. This command:
- Pulls all active skills from the remote Postgres instance.
- Compares them to your local workspace files under
.gemini/skills/and.claude/skills/. - Detects modifications, prompting you with version conflict resolutions (Push, Pull, or Merge) before modifying the database or local files.
These environment variables configure MemoryOps. They can be placed in a .env file in the working directory or injected into the container shell.
| Variable Name | Required | Default Value | Description |
|---|---|---|---|
DATABASE_URL |
Yes | postgres://memoryops:memoryops@localhost:5432/memoryops |
Connection string for the PostgreSQL database. |
REDIS_URL |
Yes | redis://localhost:6379 |
Connection string for the Redis queue. |
QDRANT_URL |
Yes | http://localhost:6334 |
gRPC/HTTP URL for Qdrant vector database. |
CONFIG_PATH |
No | config.toml |
Path to the TOML configuration file. |
APP_HOST |
No | 0.0.0.0 |
Host IP address the API server binds to. |
APP_PORT |
No | 8080 |
Port the API server listens on. |
APP_ENV |
No | development |
Setting to production enforces strict secret key validation. |
APP_SECRET_KEY |
Yes (in production) | — | Cryptographic key used to encrypt skill credentials. Must be stable across restarts. |
WORKSPACE_CREATION_SECRET |
Required only when workspace creation is enabled | — | Secret token required to authenticate workspace creation requests (x-admin-token). Rotate or remove after bootstrap. |
WORKSPACE_CREATION_ENABLED |
No | true in local compose, false in production overlay |
Set to false after initial bootstrap to disable POST /v1/workspaces. |
SKIP_MIGRATIONS |
No | false |
When set to true, bypasses database migrations on API server startup. |
TRUSTED_PROXY_CIDRS |
No | 127.0.0.1/32 |
Comma-separated CIDR blocks representing trusted reverse proxies. |
For the complete production hardening checklist, see security-production.md.
| MEMORYOPS_ALLOW_PRIVATE_IPS | No | false | Set to true to allow skills to target internal/loopback IP addresses. |
| QDRANT_CHECK_COMPATIBILITY | No | false | Bypasses major/minor version verification between client library and Qdrant database. |
| MCP_TRANSPORT | No | stdio | Transport type for the MCP server (stdio or http). |
| MCP_PORT | No | 3003 | Port for the MCP server when transport is http. |
| RUST_LOG | No | info | Logging framework level filtering (trace, debug, info, warn, error). |
| OPENAI_API_KEY | Conditional | — | Required if using OpenAI LLM or OpenAI embedding providers. |
| ANTHROPIC_API_KEY | Conditional | — | Required if using Anthropic LLM provider. |
| GEMINI_API_KEY | Conditional | — | Required if using Google Gemini LLM provider. |
| OPENROUTER_API_KEY | Conditional | — | Required if using OpenRouter LLM provider. |
| HF_API_KEY | Conditional | — | Required if using Hugging Face Inference Router LLM provider. |
| MEMORYOPS_WORKSPACE_ID | No | — | Runtime Workspace ID injected into the Frontend. |