Back to the catalog

services

Bundle OKF 0.1 · 5 conceitos · DatacationOrg/llms4eu

Open source Repository Open in the app JSON README (API)

About

# services

* [Gostilna Pohle](gostilna-pohle.md) - A restaurant and guesthouse in Brestanica, Slovenia, offering classic cuisine, accommodation, and tennis court reservations.
* [Peninoteka na Gradu Rajhenburg](peninoteka-rajhenburg.md) - A sparkling wine tasting room located within Rajhenburg Castle in Krško, Slovenia, offering guided degustations of local wines.
* [Restaurant A3](restaurant-a3.md) - A restaurant located within Rajhenburg Castle in Slovenia, offering culinary experiences inspired by the castle's history and local fish farming traditions.
* [The Chocolaterie](chocolaterie-rajhenburg.md) - A chocolate production facility at the House of Mozer near Rajhenburg Castle in Slovenia, continuing the Trappist monks' legacy of industrial chocolate-making in the region.
* [The Museum Shop](rajhenburg-castle-museum-shop.md) - A museum shop at Rajhenburg Castle in Slovenia offering hand-crafted artisan products, modern artistic goods, and museum publications thematically linked to

Details

Kind
OKF bundles
Topic
Government & public data
Publisher
datacationorg
Origin
okf_github
Category
dados
Version
0.1
Open pull requests
5
Last push
2026-09-07T15:46:35Z
Repository state
ativo
Language
Python
Added
2026-09-08 16:04:09
Updated
2026-09-08 16:04:09
Origin id
DatacationOrg/llms4eu:data/okf/tourism/services/index.md

README

# LLMs4EU - Tourism - RAG

Local RAG for tourism places. SQLite stores canonical content, and Chroma stores
derived vector indexes. Both run entirely inside the Python environment.

## Requirements

- [uv](https://docs.astral.sh/uv/)
- [just](https://just.systems/)
- [Ollama](https://ollama.com/) — only needed for `just ask`, `just scrape`,
  and `just eval-generate`

## Setup

```bash
uv sync --extra dev   # install all dependencies
cp .env.example .env  # set local paths (defaults work out of the box)
```

```bash
ollama pull gemma4:e4b
# Optional for eval question generation:
ollama pull gemma4:26b-a4b-it-q4_K_M
ollama pull gpt-oss:20b
```

## Usage

```bash
just init        # load seed data into SQLite
just index       # embed places and rebuild Chroma
just ask What place is best for a quiet forest walk near water?  # retrieve + LLM answer
just scrape https://example.com  # crawl a site, ingest into SQLite, reindex
just scrape-web  # start the local scraping UI
just eval-index qwen  # rebuild the default qwen chunk vector index
just test        # run the non-LLM test suite
```

Search defaults to top 10 final results. Override per query:

```bash
uv run python -m src.rag.search "lake picnic" --limit 5
```

Experimental retrieval eval over scraped Markdown pages:

```bash
just eval-chunks
just eval-index qwen
just eval-generate 10
just eval qwen,sparse
just eval qwen4b_rerank,qwen4b_hybrid,qwen4b_hybrid_rerank
```

## Shape

```text
data/           tracked seed fixture
sql/            one-table schema, portable to SQLite and Postgres
src/db/         SQLite initialize and place queries
src/preprocess/ rebuild derived data from SQL rows
src/indexing/   provider-shaped vector indexing
src/vector_store/  Chroma collection, upsert, vector search
src/retrieval/  chunk retrieval methods and catalog
src/rag/        place search and answer scripts
src/eval/       chunked raw-page retrieval evaluation
src/scraping/   crawler, transform, ingest, scraping UI
src/shared/     schema, embeddings, env, LLM helper
tests/          data contract, retrieval, scrape transform
```

Durable reference databases live under `data/db/` with descriptive names such
as `pages.db`. Regenerable vector cache artifacts live under
`data/cache/chroma/`. SQLite page chunks are the source of truth for chunk text;
Chroma collections are derived indexes over those chunks. Use `.env` overrides
for private scratch paths under `.local/`.

---

## Roadmap

This repo focuses on building a clean, portable place database as the foundation
for a larger RAG system.

**Current:** SQLite + Chroma, everything local, no services needed.

**Next step:** swap storage backends for larger shared runs:
- SQLite → **Supabase** (free tier, under 500 MB, shared across teams)
- Chroma → **Qdrant** (free tier, managed vector index)

The SQL schema and module boundaries are designed to make that swap small.
Before scaling up, we need to align with other teams on what data collection
tools and shared infrastructure are available.

## Docs

- [docs/architecture-decisions.md](docs/architecture-decisions.md): durable decisions and why they matter.
- [experiments/indexing/README.md](experiments/indexing/README.md): retrieval experiments and evaluation protocol.
- [docs/retrieval-results-agentic.md](docs/retrieval-results-agentic.md): current agentic retrieval report.

More