Turn 94 AI sources into a deduplicated bilingual daily digest
Self-host the engine behind inbrief.info to aggregate, score and summarize AI news from RSS, Hacker News, Reddit, X, YouTube and GitHub Trending.
- LinkLoot access
- Free
- Provider costs
- Unknown
What you get from it
What it does
This open-source project aggregates AI-related content from 94 distinct sources, including lab blogs, Hacker News, Reddit, X, YouTube transcripts and GitHub Trending. It runs every item through an LLM pipeline that filters for relevance, generates English and Chinese summaries, extracts tags and assigns an importance score. The system uses a two-stage deduplication process—SimHash fingerprinting followed by pgvector cosine similarity—to ensure the final daily digest is clean and non-repetitive before storing results in PostgreSQL.
Who it helps
Developers and researchers who want a curated, automated view of the AI landscape without manually checking dozens of feeds. It is particularly useful for teams needing bilingual (EN/ZH) insights or those who prefer self-hosting to control their data pipeline. If you only want to read the output, the hosted site at inbrief.info is available, but this repo provides the full engine for customization.
Getting started
The repository includes a pipeline_runner that orchestrates fetchers, content processing, and AI curation. To begin, set up a PostgreSQL database; the schema auto-creates on first run. You must configure an OpenAI-compatible API endpoint for the LLM steps (DeepSeek is the default). Use admin_rss.py to add your preferred sources, as the engine ships with an empty source table despite documenting 94 production examples. Anti-bot measures are layered, using curl_cffi, DrissionPage, and Playwright with persistent profiles for platforms like Reddit and X.
Limits and costs
While the software is MIT-licensed, self-hosting requires infrastructure resources: a server for the Python application, a PostgreSQL instance, and potentially headless browser environments for scraping protected sites. LLM usage incurs token costs depending on your provider. Optional Aliyun Green text moderation can be enabled via keys, though it is skipped if absent. The complexity of maintaining scraper resilience against platform changes is significant.
Source links
Discussion
Share practical experience, questions, or warnings with the community.
Sign in to join the discussion and vote on comments.
Sign in