Enterprise RAG Platform — MEDZ
A web platform developed during a two-month internship: a client-facing interface, a PostgreSQL data layer, geographic visualization, and two AI assistants — technical and navigation — backed by a retrieval-augmented generation pipeline with bi-encoder retrieval, cross-encoder reranking and Gemma 2.
- Context
- Internship · 2 months
- Role
- Individual implementation under professional supervision
- Timeline
- July 2026 — September 2026
- Stack
- Next.js · React · PostgreSQL · pgvector · Prisma · LangChain · BGE-M3 · Cross-encoder reranking · Gemma 2 · Ollama
01 — Overview
Overview
During a two-month internship, I redesigned and developed MEDZ's web platform and integrated two AI assistants into it. The work was an individual implementation carried out under professional supervision.
- Professional, client-facing web interface.
- PostgreSQL database managed with Prisma ORM.
- Mapping / geographic visualization.
- A technical AI assistant and a navigation assistant built on retrieval-augmented generation.
MED-Z works in Design, development, commercialization, and management of business and industrial parks
02 — Problem
Problem
after long discussions with the company we identified some weaknesses in the existing platform and user needs for the AI assistants.
03 — Architecture
Architecture
- User query
- EmbeddingBGE-M3 (bi-encoder)
- PostgreSQL + pgvectorSemantic retrieval
- Top 25 candidates
- Cross-encoder reranking
- Top 5 contexts
- Gemma 2Served with Ollama
- Generated answer
Retrieval runs in two stages. A bi-encoder (BGE-M3) embeds the query and retrieves a broad candidate set from PostgreSQL with pgvector. A cross-encoder then rescores those candidates jointly with the query, and only the highest-ranked contexts are passed to Gemma 2 for generation.
04 — Data / Inputs
Data / Inputs
Document used in chunking was generated manually that explain the insights and processes involved and also document that explain the navigation and guidance of using the web-plateforme
05 — Methodology
Methodology
- 01EmbedQueries and documents are embedded with BGE-M3.
- 02RetrieveSemantic search over pgvector returns the top 25 candidate chunks.
- 03RerankA cross-encoder rescores the candidates against the query and keeps the top 5.
- 04GenerateGemma 2, served through Ollama, generates the answer from the selected contexts. LangChain wires the pipeline together.
The two-stage design trades a small amount of latency for precision: the bi-encoder is fast enough to search the whole index, while the cross-encoder is more accurate but only practical on a short list.
the bi-encoder used is BGE-M3 using cosine similarity and the cross-encoder was BERT.
06 — Engineering Implementation
Engineering Implementation
- Next.js and React for the web platform and assistant interfaces.
- PostgreSQL with pgvector as a single store for application data and embeddings.
- Prisma ORM for schema and data access.
- LangChain for the retrieval and generation pipeline.
- Ollama for serving Gemma 2.
Mapping library type is Object-Relational Mapping (ORM) and the specific library used is Prisma.
07 — Evaluation / Results
Evaluation / Results
08 — Challenges & Trade-offs
Challenges & Trade-offs
the real challenges was to implement the solution based on both axes AI and development web .
09 — Resources
Resources
https://github.com/charfx/Medz_project.


