Finkkle / Architecture

Architecture

Crawling & Indexing

Learn how Finkkle discovers web content, populates the search tables, and processes index calculations.

Indexer reference · Finkkle documentation

Finkkle crawl eligibility

Arachnid uses these requirements as a quality floor before it follows links or sends a page to the Finkkle index. Being crawled is not a guarantee of ranking or inclusion; pages may still be excluded by later relevance, safety, duplication, or abuse systems.

Technical requirements

  • Serve a public HTML page over HTTPS with a successful 2xx response. Redirects should resolve to the preferred canonical URL.
  • Do not block Arachnid in robots.txt, HTTP X-Robots-Tag, or a robots meta tag. noindex means the page is excluded; nofollow prevents its links being used as a crawl frontier.
  • Keep the page available without a login, challenge, paywall, or user interaction. Authentication, checkout, account, API, and private application routes are not search pages.
  • Use a stable canonical URL, a meaningful <title>, a useful meta description, a declared language, and a clear heading structure. These are quality signals and help Finkkle explain the result.
  • Provide meaningful visible content in the HTML response. JavaScript-only shells, empty templates, placeholder pages, and pages with only navigation or advertisements may be excluded.

Content quality floor

  • Write for a real reader: original, specific, useful information with enough context to stand on its own.
  • Avoid copied or near-duplicate pages, doorway pages, automatically generated query/filter combinations, keyword stuffing, spun text, lorem ipsum, and low-effort AI or template output.
  • Do not cloak content, hide text, imitate another site, distribute malware/phishing, or use misleading titles and structured data.
  • Use descriptive internal links, accurate image alt text, and structured data only when it describes the visible page. Avoid intrusive interstitials and excessive ads.
  • For large sites, publish an XML sitemap with canonical URLs and accurate lastmod values. Keep response times, page size, redirects, and crawl-rate impact reasonable.

When a page is fetched but fails the quality floor, Arachnid records the exclusion reason in Console, does not index the page, and does not expand the crawl from that page. Fix the page and run a new preflight. A preflight must crawl at least one eligible page before a full crawl can be started.

Storage Models

Finkkle utilizes a SQL schema optimized for fast FTS text querying and vector queries. Two environments are supported:

  • Local SQLite (search.db): Used for local prototyping, testing, and single-node Node.js executions.
  • Turso / libSQL Database: Remote, replicated serverless database used during production server operations and Vercel serverless deployments.

Text indices are maintained in a pages_fts virtual table to execute rapid prefix, token, and proximity lookups.

The Crawl Stack

Finkkle includes two separate crawler setups:

Classic Crawler

Found under src/crawler/. Provides straightforward, sequential scraping loops. Reads target sitemaps and crawls static HTML documents, respecting standard robots.txt limitations.

Phase 1 Crawler

Located in src/crawler/phase1/. Designed as a distributed crawl coordinator and worker system. It leverages playbooks, sitemaps, worker queues, and runs Playwright browser instances to capture dynamic JavaScript-rendered pages.

Crawl commands include:

  • npm run crawl: Executes classic crawler.
  • npm run crawl:phase1: Runs the modern Phase 1 crawler.
  • npm run crawl-coordinator: Launches the queue coordinator.
  • npm run crawl-worker: Starts a crawl worker process.

Index Maintenance Scripts

Following ingestion, several scripts must run to update relevance scores, authority ranks, and embedding data:

PageRank

Calculates document relationship linkages. Run with npm run pagerank.

Domain Authority

Assigns score factors based on domain host credibility. Run with npm run domain-authority.

Vector Embeddings

Creates vector snapshots of crawled text segments. Run with npm run embed.

Post-Crawl Pipeline

Executes PageRank, authority scoring, and embedding generations sequentially. Run with npm run post-crawl.