💡 Feature Description
As a platform operator running vector-capable databases, I want a
native embedding service in the ai-services stack so that I can
generate vectors locally on IBM Power without depending on a remote
endpoint or managing a custom container outside the catalog.
Currently there is no packaged embedding generation service in the
ai-services stack. Teams using vector databases must operate a
separate embedding service themselves or call a remote endpoint.
This feature adds embedding generation as a first-class ai-service,
consistent with how other services (summarize, similarity, chatbot)
are deployed and managed today.
The service exposes an OpenAI-compatible POST /v1/embeddings
endpoint. It runs CPU-only on IBM Power11 ppc64le (RHEL 9, no Spyre
cards required), deployable via the ai-services catalog UI using
Podman or as a standalone Linux sidecar container alongside any
application stack. Any vector database framework that consumes an
OpenAI-compatible embeddings endpoint can use this service.
❗ Why This Matters
Teams using vector databases (OpenSearch, pgvector, Oracle, Milvus,
etc.) need an embedding service to generate vectors from text or
images. Without a native option in the stack, every team must:
- Call a remote endpoint (latency, data leaves the system), or
- Manage a custom container outside the ai-services catalog
Embedding generation is a core primitive for any RAG or vector search
workflow. Adding it as a native service means:
- Any vector database on Power can use it through a standard API
- No Spyre cards required — works on any Power11 LPAR
- Consistent deployment model with the rest of the ai-services catalog
- Can be offered as a standalone service or bundled in architectures
🧩 Requirements
- Expose
POST /v1/embeddings compatible with the OpenAI embeddings
API (text and image inputs)
- Support models validated by DMG and compatible with the vLLM runtime
(e.g. clip-vit-base-patch32, siglip-base-patch16-224)
- Model selection via
MODEL_NAME environment variable
- Run CPU-only on IBM Power11 ppc64le RHEL 9 — no Spyre cards required
- Deployable via Podman through the ai-services catalog UI
- Return float32 vectors consumable by any vector database framework
- Include unit tests with mocked model inference (no hardware required
to run CI)
- Pass existing CI workflows:
go.yaml and service-unit-tests.yml
📚 Additional Details
- Jira: AISERVICES-1811
- Reference implementation tested on IBM Power11 LPAR, RHEL 9.6,
Podman runtime
- Fork with working implementation:
https://github.com/nava-dba/embedding-services
- Release binary tested:
v0.4.11-embedding (ai-services-linux-ppc64le)
- No new global infrastructure required — service is self-contained,
no OpenSearch or PostgreSQL dependency
💡 Feature Description
As a platform operator running vector-capable databases, I want a
native embedding service in the ai-services stack so that I can
generate vectors locally on IBM Power without depending on a remote
endpoint or managing a custom container outside the catalog.
Currently there is no packaged embedding generation service in the
ai-services stack. Teams using vector databases must operate a
separate embedding service themselves or call a remote endpoint.
This feature adds embedding generation as a first-class ai-service,
consistent with how other services (summarize, similarity, chatbot)
are deployed and managed today.
The service exposes an OpenAI-compatible
POST /v1/embeddingsendpoint. It runs CPU-only on IBM Power11 ppc64le (RHEL 9, no Spyre
cards required), deployable via the ai-services catalog UI using
Podman or as a standalone Linux sidecar container alongside any
application stack. Any vector database framework that consumes an
OpenAI-compatible embeddings endpoint can use this service.
❗ Why This Matters
Teams using vector databases (OpenSearch, pgvector, Oracle, Milvus,
etc.) need an embedding service to generate vectors from text or
images. Without a native option in the stack, every team must:
Embedding generation is a core primitive for any RAG or vector search
workflow. Adding it as a native service means:
🧩 Requirements
POST /v1/embeddingscompatible with the OpenAI embeddingsAPI (text and image inputs)
(e.g. clip-vit-base-patch32, siglip-base-patch16-224)
MODEL_NAMEenvironment variableto run CI)
go.yamlandservice-unit-tests.yml📚 Additional Details
Podman runtime
https://github.com/nava-dba/embedding-services
v0.4.11-embedding(ai-services-linux-ppc64le)no OpenSearch or PostgreSQL dependency