# Setup Guide This document walks through installing, configuring, and running the web scraper — including the optional Qdrant vector search and MCP server. ## Requirements - Node.js 16 or newer - npm - SQLite (ships via `better-sqlite3`, no external server needed) - Linux with systemd (only if you want to install it as a service) - Optionally: a Qdrant instance and an OpenRouter API key for semantic search ## Install ```bash git clone cd webshop-scraper npm install ``` The `npm install` step will also install dependencies for the web UI under `web/`. ## Configure Copy `.env.example` to `.env` and fill in values: ```bash cp .env.example .env ``` ### Environment variables All environment variables read by the code live in `src/config/index.ts`. Nothing else is consulted. | Variable | Required | Default | Purpose | |---|---|---|---| | `API_KEY` | **yes** | — | Bearer token required for every `/api/*` endpoint. Startup fails without it. | | `HOST` | no | `0.0.0.0` | HTTP bind address. Use `127.0.0.1` to only accept local connections. | | `PORT` | no | `3000` | HTTP port. | | `MAX_CONCURRENT_JOBS` | no | `3` | Parallel scraping worker slots. | | `LOG_HTTP_VERBOSE` | no | `false` | When `true`, HTTP access logs include user agent and content length. | | `QDRANT_ENABLED` | no | `false` | Master switch for the Qdrant integration. Must be `true` for embeddings and MCP search to work. | | `QDRANT_API_URL` | no | `http://localhost:6333` | Qdrant HTTP endpoint. | | `QDRANT_API_KEY` | no | *(empty)* | Qdrant API key; leave empty for unauthenticated local instances. | | `QDRANT_DELETION_DELAY_HOURS` | no | `24` | Grace period before queued Qdrant deletions actually run. | | `OPENROUTER_API_KEY` | no | *(empty)* | Required if `QDRANT_ENABLED=true`; used to generate embeddings. | | `MCP_ENABLED` | no | `false` | Mount the MCP Streamable HTTP endpoints at `/mcp`. Also requires `QDRANT_ENABLED=true`. | | `MCP_PORT` | no | `3001` | Reserved port (the MCP server currently mounts onto the main HTTP server under `/mcp`). | Example minimal `.env`: ```bash API_KEY=change-me HOST=0.0.0.0 PORT=3000 MAX_CONCURRENT_JOBS=3 ``` Example `.env` with Qdrant + MCP: ```bash API_KEY=change-me PORT=3000 QDRANT_ENABLED=true QDRANT_API_URL=http://localhost:6333 QDRANT_API_KEY= QDRANT_DELETION_DELAY_HOURS=24 OPENROUTER_API_KEY=sk-or-v1-... MCP_ENABLED=true ``` ## Site types The scraper supports two site types, set when a job is created or updated via `PATCH /api/sites/:id/settings`: - **`webshop`** — e-commerce sites built on ShopRenter, WooCommerce, or Shopify. The scraper uses the platform's sitemap to discover pages and classifies them into content categories. - **`website`** — any generic website. When no sitemap is found, the scraper falls back to the link discovery engine (see below). When you create a job with `POST /api/jobs`, the system auto-detects the site type and platform via `POST /api/sites/detect`. You can override the detected values with `PATCH /api/sites/:id/settings`. ## Link discovery When a site has no sitemap (common for generic websites), the scraper uses a BFS-based link discovery engine that crawls same-origin links starting from the site's base URL. Two settings control the crawl: | Setting | Default | Description | |---|---|---| | `max_crawl_depth` | `3` | Maximum link-follow depth from the start URL. | | `max_pages` | `200` | Maximum number of pages to visit. | Both can be configured per site via `PATCH /api/sites/:id/settings`. Discovered URLs are then classified into content categories the same way sitemap URLs are. You can also add **URL exclusion patterns** (`POST /api/sites/:siteId/url-exclusions`) using glob syntax (e.g. `*/admin/*`, `*/wp-json/*`) to skip unwanted URLs during scraping. ## Build ```bash npm run build ``` This compiles TypeScript (`tsc`) and then runs `npm run build:web` to build the bundled web UI under `web/`. The web UI is served at `/ui` from the main HTTP server. ## Run Ask the user before starting the server — don't start it yourself. ```bash # Production npm start # Development (ts-node, no build required) npm run dev # TypeScript watch mode npm run watch ``` Once up: - API: `http://:` - Web UI: `http://:/ui` - Live API docs: `http://:/doc` - OpenAPI spec: `http://:/doc/openapi` ## Install as a systemd service ```bash sudo ./scripts/install.sh ``` ## Reset and restore Both scripts live in `scripts/` and create timestamped backups under `/var/backups/webshop-scraper/`. ```bash sudo ./scripts/reset.sh # wipe back to a clean state (backs up first) sudo ./scripts/restore.sh # restore from a previous backup ``` A backup contains: - `data/shops.db` (SQLite database) - `.env` - Service logs from `journalctl` ## Qdrant vector search (optional) The scraper can embed scraped content and push it to a Qdrant collection so downstream systems (or the MCP tools) can do semantic search over it. Everything is gated on `QDRANT_ENABLED=true`. ### 1. Run Qdrant ```bash # Local Docker instance with persistent storage docker run -p 6333:6333 -v $(pwd)/qdrant_storage:/qdrant/storage qdrant/qdrant ``` Or use Qdrant Cloud and point `QDRANT_API_URL` / `QDRANT_API_KEY` at your cluster. ### 2. Get an OpenRouter API key The `EmbeddingService` (`src/services/EmbeddingService.ts`) calls OpenRouter through the `openai` client. Sign up at , create a key, and put it in `OPENROUTER_API_KEY`. ### 3. Turn on the integration Set `QDRANT_ENABLED=true` in `.env` and restart the server. ### 4. Enable Qdrant per site Embedding happens per site — you opt each site in. 1. Make sure the site has a `custom_id` set (`PATCH /api/sites/:id/custom-id`). Once Qdrant is enabled for a site, its `custom_id` becomes immutable — this prevents orphaned collections. 2. Toggle Qdrant on via either the web UI or the API: ```bash curl -X PUT http://localhost:3000/api/sites//qdrant/toggle \ -H "Authorization: Bearer $API_KEY" \ -H "Content-Type: application/json" \ -d '{"enabled": true}' ``` 3. The next successful scrape of that site triggers embedding generation via `QdrantEmbeddingWorkflow`. ### 5. Collection layout - One collection per `{custom_id}-{category}`, e.g. `shop123-shipping`, `shop123-faq`. - Categories: `shipping`, `contacts`, `terms`, `faq`, `services`, `about`, `team`, `blog`, `pricing`, `testimonials`, `gallery`, `landing`, `other`. - Each page is stored as one or more chunks (via `ChunkingService`); point IDs are deterministic: `SHA256(url + content_hash + chunk_index)`. Same URL + same content ⇒ same point ID, so re-embedding is a no-op; content changes yield a fresh point. Verified in `src/services/QdrantEmbeddingWorkflow.ts` and `src/services/QdrantService.ts`. ### 6. Legacy → current migration If you have data from a previous version that used per-URL collections (`{custom_id}-{category}_{url_hash}`), migrate using the per-site endpoints: ```bash # 1. See what would happen curl -H "Authorization: Bearer $API_KEY" \ http://localhost:3000/api/sites//qdrant/migration-status # 2. Copy points into the new consolidated collections curl -X POST -H "Authorization: Bearer $API_KEY" \ http://localhost:3000/api/sites//qdrant/migrate # 3. Dry-run cleanup of old collections curl -X POST -H "Authorization: Bearer $API_KEY" \ -H "Content-Type: application/json" \ -d '{"dryRun": true}' \ http://localhost:3000/api/sites//qdrant/cleanup-old-collections # 4. Actually delete the old collections curl -X POST -H "Authorization: Bearer $API_KEY" \ -H "Content-Type: application/json" \ -d '{"dryRun": false}' \ http://localhost:3000/api/sites//qdrant/cleanup-old-collections ``` ## MCP (Model Context Protocol) When `MCP_ENABLED=true` **and** `QDRANT_ENABLED=true`, `src/api/server.ts` mounts the MCP Streamable HTTP transport at `/mcp`. 14 tools are exposed (1 listing + 13 search tools, one per content category); see `docs/API.md` for the list and the shapes. Point any MCP client at: ``` http://:/mcp ``` The server uses the stateless mode of `StreamableHTTPServerTransport`, so each request is self-contained. Quick health check: ```bash curl http://localhost:3000/mcp/health ``` ## Troubleshooting | Symptom | Where to look | |---|---| | `API_KEY environment variable is required` at startup | `.env` missing or not loaded; `API_KEY` must be set. | | `401` on every request | Missing `Authorization: Bearer ...` header, wrong token, or using a different header (the server only accepts Bearer). | | Site exists but MCP tools return "Search not available" | Either `QDRANT_ENABLED=false` or the site's `qdrant_enabled` flag is false or the site has no `custom_id`. Check `GET /api/sites/:siteId/qdrant`. | | `qdrant_errors` piling up | Inspect via `GET /api/sites/:siteId/qdrant/errors`; the likely culprits are OpenRouter quota, wrong `OPENROUTER_API_KEY`, or unreachable `QDRANT_API_URL`. | | Embeddings are not being created after scrape | Check the site is toggled on (`PUT /api/sites/:siteId/qdrant/toggle`) and re-run with `POST /api/sites/:siteId/qdrant/re-embed` or `/sync`. | For logs: ```bash # systemd sudo journalctl -u webshop-scraper -f # development npm run dev ``` ## Database SQLite file at `data/shops.db`. Schema lives in `src/database/SiteDatabase.ts` and is applied on startup — no manual migration step. Qdrant-related tables (`shop_embeddings`, `qdrant_deletion_queue`, `qdrant_errors`, `mcp_logs`) are created automatically whether or not Qdrant is enabled.