Turn 94 AI sources into a deduplicated bilingual daily digest

Self-host the engine behind inbrief.info to aggregate, score and summarize AI news from RSS, Hacker News, Reddit, X, YouTube and GitHub Trending.

LinkLoot access
Free
Provider costs
Unknown
The useful part2 min read

What you get from it

What it does

This open-source project aggregates AI-related content from 94 distinct sources, including lab blogs, Hacker News, Reddit, X, YouTube transcripts and GitHub Trending. It runs every item through an LLM pipeline that filters for relevance, generates English and Chinese summaries, extracts tags and assigns an importance score. The system uses a two-stage deduplication process—SimHash fingerprinting followed by pgvector cosine similarity—to ensure the final daily digest is clean and non-repetitive before storing results in PostgreSQL.

Who it helps

Developers and researchers who want a curated, automated view of the AI landscape without manually checking dozens of feeds. It is particularly useful for teams needing bilingual (EN/ZH) insights or those who prefer self-hosting to control their data pipeline. If you only want to read the output, the hosted site at inbrief.info is available, but this repo provides the full engine for customization.

Getting started

The repository includes a pipeline_runner that orchestrates fetchers, content processing, and AI curation. To begin, set up a PostgreSQL database; the schema auto-creates on first run. You must configure an OpenAI-compatible API endpoint for the LLM steps (DeepSeek is the default). Use admin_rss.py to add your preferred sources, as the engine ships with an empty source table despite documenting 94 production examples. Anti-bot measures are layered, using curl_cffi, DrissionPage, and Playwright with persistent profiles for platforms like Reddit and X.

Limits and costs

While the software is MIT-licensed, self-hosting requires infrastructure resources: a server for the Python application, a PostgreSQL instance, and potentially headless browser environments for scraping protected sites. LLM usage incurs token costs depending on your provider. Optional Aliyun Green text moderation can be enabled via keys, though it is skipped if absent. The complexity of maintaining scraper resilience against platform changes is significant.

Community

Discussion

Share practical experience, questions, or warnings with the community.

0

Sign in to join the discussion and vote on comments.

No comments yet. Start the discussion.
Keep exploring

More from this topic

More in Tools & Apps