{
  "markdown": "<div align=\"center\">\n\n# 🔍 ask-search\n\n**Self-hosted web search for AI agents — zero API key, full privacy**\n\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)\n[![OpenClaw](https://img.shields.io/badge/OpenClaw-Skill-blue.svg)](https://github.com/openclaw/openclaw)\n[![Python](https://img.shields.io/badge/Python-3.8%2B-blue.svg)](https://python.org)\n\n*Works in OpenClaw · Claude Code · Antigravity · Any CLI*\n\n</div>\n\n---\n\n## 😡 The Problem\n\nYour AI agent wants to search the web, but:\n- Brave Search API: $3/1000 queries, rate limited\n- Google Custom Search: $5/1000, daily caps\n- Bing API: paid, complex setup\n- Built-in web search: sends queries to third-party servers\n\n**You just want your local agent to search Google without paying or leaking queries.**\n\n## ✅ The Solution\n\n`ask-search` wraps [SearxNG](https://github.com/searxng/searxng) — a self-hosted meta search engine that aggregates Google, Bing, DuckDuckGo, Brave and 70+ more sources. One command, all results, zero cost.\n\n```bash\nask-search \"Claude Code vs Cursor 2026\"\n```\n\n```\n[1] Claude Code Is Eating Cursor's Lunch\n    https://techcrunch.com/2026/...\n    After Anthropic launched Claude Code with agent mode...\n    [google,brave]\n\n[2] Why I switched from Cursor to Claude Code\n    https://reddit.com/r/LocalLLaMA/...\n    ...\n```\n\n## 📊 Compatibility\n\n| Environment | Integration | Status |\n|-------------|-------------|--------|\n| **OpenClaw** | CLI Skill (`SKILL.md`) | ✅ |\n| **Claude Code** | CLI command | ✅ |\n| **Antigravity** | MCP Server | ✅ |\n| Any shell | `ask-search` CLI | ✅ |\n\n## 🚀 Quick Start\n\n### 30-second version (if SearxNG already running)\n\n```bash\ngit clone https://github.com/ythx-101/ask-search\ncd ask-search\nbash install.sh\nask-search \"hello world\"\n```\n\n### Full setup with Docker Compose (Recommended)\n\n```bash\n# Clone and navigate to the project\ngit clone https://github.com/ythx-101/ask-search\ncd ask-search\n\n# Generate a secure secret key\npython3 -c \"import secrets; print('SEARXNG_SECRET=' + secrets.token_hex(32))\"\n\n# Copy .env.example to .env and set your secret key\ncp searxng/.env.example searxng/.env\n# Edit .env with your generated secret\n\n# Start SearxNG\ncd searxng && docker-compose up -d\n\n# Configure ask-search with your secret\nexport SEARXNG_SECRET=\"your-generated-secret\"\nask-search \"hello world\"\n```\n\n### Full setup (SearxNG + ask-search)\n\n**Step 1: Deploy SearxNG**\n\n```bash\n# Docker (recommended)\ndocker run -d --name searxng \\\n  -p 127.0.0.1:8080:8080 \\\n  -e SEARXNG_SECRET_KEY=your-secret-key \\\n  searxng/searxng\n\n# Or Docker Compose — see searxng/docker-compose.yml in this repo\n```\n\n**Step 2: Enable JSON output**\n\nEdit SearxNG `settings.yml`:\n```yaml\nsearch:\n  formats:\n    - html\n    - json\n```\n\n**Step 3: Install ask-search**\n\n```bash\nbash install.sh\n```\n\n**Step 4: Use it**\n\n```bash\nask-search \"your query\"\n```\n\n## 🔧 Usage\n\n```bash\nask-search \"query\"                    # top 10 results\nask-search \"query\" --num 5            # limit results\nask-search \"AI news\" --categories news # news only\nask-search \"query\" --lang zh-CN       # language filter\nask-search \"query\" --urls-only        # URLs only (pipe to web_fetch)\nask-search \"query\" --json             # raw JSON\nask-search \"query\" -e google,brave    # specific engines\n```\n\n## ⚙️ Configuration\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `SEARXNG_URL` | `http://localhost:8080` | SearxNG endpoint |\n| `SEARCH_PROVIDER` | `searxng` | Search backend: `searxng` or `tavily` |\n| `TAVILY_API_KEY` | *(none)* | Tavily API key (required when `SEARCH_PROVIDER=tavily`) — get one at [app.tavily.com](https://app.tavily.com) |\n\n```bash\n# SearxNG (default)\nexport SEARXNG_URL=\"http://localhost:8080\"\nask-search \"query\"\n\n# Tavily (opt-in)\nexport SEARCH_PROVIDER=tavily\nexport TAVILY_API_KEY=tvly-YOUR_API_KEY\nask-search \"query\"\n```\n\n## 🤖 MCP Integration (Antigravity / Claude Code)\n\n```json\n{\n  \"mcpServers\": {\n    \"ask-search\": {\n      \"command\": \"python3\",\n      \"args\": [\"/path/to/ask-search/mcp/server.py\"],\n      \"env\": {\n        \"SEARXNG_URL\": \"http://localhost:8080\",\n        \"TAVILY_API_KEY\": \"tvly-YOUR_API_KEY\"\n      }\n    }\n  }\n}\n```\n\nRequires: `pip install mcp`\n\nMCP tools provided:\n- `web_search` — search via SearxNG\n- `web_search_news` — news search via SearxNG\n- `web_search_tavily` — search via Tavily API (requires `TAVILY_API_KEY`)\n\n## 🦞 OpenClaw Skill\n\nCopy `SKILL.md` to your OpenClaw skills directory, or:\n\n```bash\n# OpenClaw skill loader\nskill-add https://github.com/ythx-101/ask-search\n```\n\nThen in OpenClaw:\n```\nask-search \"latest news about X\"\n```\n\n## 🤝 Agent Workflow\n\n```bash\n# 1. Search\nask-search \"React Server Components performance 2026\" --num 10\n\n# 2. Got a promising URL? Deep-dive:\n# Pass the URL to web_fetch / curl for full content\n\n# 3. News mode\nask-search \"GPT-5 release\" --categories news --lang en\n```\n\n## 📁 Structure\n\n```\nask-search/\n├── scripts/\n│   └── core.py        # Main logic, CLI entry point\n├── mcp/\n│   └── server.py      # MCP server for AG/CC integration\n├── install.sh         # Installer\n├── SKILL.md           # OpenClaw skill descriptor\n└── README.md\n```\n\n## ⚠️ SearxNG Setup Notes\n\n- Must enable `json` in `search.formats` in `settings.yml`\n- Bot detection may block requests from some IPs — add your server IP to `pass_ip` in `limiter.toml`\n- Bind to `127.0.0.1` for security unless you need remote access\n\n## 🌐 Deep-Dive Limitations & Workarounds\n\n`ask-search` returns URLs + snippets from search engine indexes. When your agent needs the **full page content** (deep-dive via `curl` / `web_fetch`), some sites will block depending on your server's network environment:\n\n### What works out of the box\n\n| Site | Search (SearxNG) | Deep-dive (curl/fetch) | Why |\n|------|:-:|:-:|-----|\n| Most sites | ✅ | ✅ | No aggressive anti-bot |\n| Reddit | ✅ | ❌ VPS IP blocked | Reddit blocks datacenter IPs |\n| Zhihu (知乎) | ✅ | ❌ Login wall + fingerprint | Requires browser JS + login |\n| Medium | ✅ | ⚠️ Paywall | Partial content only |\n\n**Key insight**: Search always works because SearxNG queries search engines (Google, Brave, etc.), not the target sites directly. The search engines have already indexed the content. The problem only appears when your agent tries to fetch the full page.\n\n### Solution 1: SOCKS proxy via residential IP\n\nIf you have a machine on a residential network (home server, laptop, etc.), create an SSH SOCKS tunnel:\n\n```bash\n# On your VPS/server:\nssh -f -N -D 127.0.0.1:1082 user@your-home-machine\n\n# Then fetch through the proxy:\ncurl -x socks5h://127.0.0.1:1082 \"https://reddit.com/r/example/comments/xxx.json\"\n```\n\nFor Reddit specifically, append `.json` to any post URL for structured data:\n```bash\n# Returns full post + all comments as JSON\ncurl -x socks5h://127.0.0.1:1082 \\\n  \"https://www.reddit.com/r/LocalLLaMA/comments/xxxxx/post_title.json\"\n```\n\nTo persist the tunnel as a systemd service:\n```ini\n# /etc/systemd/system/socks-proxy.service\n[Unit]\nDescription=SSH SOCKS Proxy for web scraping\nAfter=network.target\n\n[Service]\nType=simple\nExecStart=/usr/bin/ssh -N -D 127.0.0.1:1082 -o ServerAliveInterval=30 -o ServerAliveCountMax=3 -o ExitOnForwardFailure=yes user@your-home-machine\nRestart=always\nRestartSec=10\n\n[Install]\nWantedBy=multi-user.target\n```\n\n### Solution 2: Headless browser for JS-heavy sites\n\nSites like Zhihu require full browser rendering. Use Playwright with the SOCKS proxy:\n\n```python\nfrom playwright.sync_api import sync_playwright\n\nwith sync_playwright() as p:\n    browser = p.chromium.launch(\n        headless=True,\n        proxy={\"server\": \"socks5://127.0.0.1:1082\"}\n    )\n    page = browser.new_page()\n    page.goto(\"https://example.com/article\")\n    content = page.inner_text(\"article\")\n```\n\n> **Note**: Some sites (Zhihu, etc.) detect headless browsers even through proxies. For these, you may need a real browser session with login cookies, or delegate the fetch to an agent running on a local machine (e.g., Claude Code on a Mac).\n\n### Solution 3: Leverage archive caches\n\nWhen direct access fails, try cached versions:\n\n```bash\n# Archive.org (works surprisingly often for Reddit)\ncurl \"https://web.archive.org/web/2026/https://reddit.com/r/example/comments/xxx\"\n\n# Google Cache (may redirect — not always reliable)\ncurl \"https://webcache.googleusercontent.com/search?q=cache:example.com/page\"\n```\n\n### Solution 4: Multi-node agent architecture\n\nIf you run agents on multiple machines (e.g., OpenClaw on VPS + Claude Code on local Mac):\n\n```\nVPS agent: ask-search \"query\" → gets URLs + snippets\n                ↓\nLocal agent: web_fetch(url) → full content (residential IP, no blocks)\n                ↓\nVPS agent: receives full text, analyzes, responds\n```\n\nThis is the most robust approach — search on your server, deep-dive from a local machine where anti-bot measures don't apply.\n\n### TL;DR\n\n| Problem | Fix |\n|---------|-----|\n| Reddit blocks your IP | SSH SOCKS proxy + `.json` API |\n| Site needs JS rendering | Playwright + proxy |\n| Site needs login (Zhihu) | Delegate to local agent or use logged-in browser |\n| Everything blocked | Fall back to search snippets + archive caches |\n\n## 🤝 Contributing\n\nIssues and PRs welcome.\n\n## Acknowledgements\n\nInspired by [Perplexica](https://github.com/ItzCrazyKns/Perplexica) —\nthe idea of using self-hosted SearxNG as a private search backend comes from there.\nask-search strips it down to a single CLI command for agent use.\n\n## 📄 License\n\n[MIT](LICENSE)\n",
  "bytes": 9457,
  "sha": "b675f7baa9c1c93034c8299527933864ab851dbd1fa1ff7e87b1960a462e7bb1",
  "repo_slug": "ythx-101/ask-search",
  "fonte": "repo",
  "truncated": false,
  "api": "https://api.agentalog.com/api/listings/skl_ythx_101_ask_search_ask_search_c69ef6ea/readme"
}