Skip to content

feat: add multimodal embedding service for IBM Power (ppc64le, CPU-only) #1467

Description

@nava-dba

💡 Feature Description

As a platform operator running vector-capable databases, I want a
native embedding service in the ai-services stack so that I can
generate vectors locally on IBM Power without depending on a remote
endpoint or managing a custom container outside the catalog.

Currently there is no packaged embedding generation service in the
ai-services stack. Teams using vector databases must operate a
separate embedding service themselves or call a remote endpoint.
This feature adds embedding generation as a first-class ai-service,
consistent with how other services (summarize, similarity, chatbot)
are deployed and managed today.

The service exposes an OpenAI-compatible POST /v1/embeddings
endpoint. It runs CPU-only on IBM Power11 ppc64le (RHEL 9, no Spyre
cards required), deployable via the ai-services catalog UI using
Podman or as a standalone Linux sidecar container alongside any
application stack. Any vector database framework that consumes an
OpenAI-compatible embeddings endpoint can use this service.

❗ Why This Matters

Teams using vector databases (OpenSearch, pgvector, Oracle, Milvus,
etc.) need an embedding service to generate vectors from text or
images. Without a native option in the stack, every team must:

  • Call a remote endpoint (latency, data leaves the system), or
  • Manage a custom container outside the ai-services catalog

Embedding generation is a core primitive for any RAG or vector search
workflow. Adding it as a native service means:

  • Any vector database on Power can use it through a standard API
  • No Spyre cards required — works on any Power11 LPAR
  • Consistent deployment model with the rest of the ai-services catalog
  • Can be offered as a standalone service or bundled in architectures

🧩 Requirements

  • Expose POST /v1/embeddings compatible with the OpenAI embeddings
    API (text and image inputs)
  • Support models validated by DMG and compatible with the vLLM runtime
    (e.g. clip-vit-base-patch32, siglip-base-patch16-224)
  • Model selection via MODEL_NAME environment variable
  • Run CPU-only on IBM Power11 ppc64le RHEL 9 — no Spyre cards required
  • Deployable via Podman through the ai-services catalog UI
  • Return float32 vectors consumable by any vector database framework
  • Include unit tests with mocked model inference (no hardware required
    to run CI)
  • Pass existing CI workflows: go.yaml and service-unit-tests.yml

📚 Additional Details

  • Jira: AISERVICES-1811
  • Reference implementation tested on IBM Power11 LPAR, RHEL 9.6,
    Podman runtime
  • Fork with working implementation:
    https://github.com/nava-dba/embedding-services
  • Release binary tested: v0.4.11-embedding (ai-services-linux-ppc64le)
  • No new global infrastructure required — service is self-contained,
    no OpenSearch or PostgreSQL dependency

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions