RSS Feed Aggregator
Python engine for reading RSS and Atom feeds and normalizing articles into a common newest-first model.
Python · feedparser · BeautifulSoup · RSS · Atom
README.md
# RSS Feed Aggregator
A small, framework-independent Python engine for pulling articles from curated RSS/Atom feeds, normalizing inconsistent feed metadata into one common model, and returning a combined newest-first collection.
The project began as a standalone NASA feed experiment and later evolved as a proper aggregator. This repository now contains the generic aggregation logic.
## Features
- Aggregate multiple RSS/Atom publications through one reusable pipeline
- Normalize feeds into a common `Article` dataclass
- Curated feed configuration with categories and fallback authors
- Publication-date parsing with updated-date fallback
- HTML cleanup and whitespace normalization for summaries
- Configurable summary truncation
- Multiple image extraction strategies:
- Media RSS content
- Media RSS thumbnails
- embedded content images
- summary images
- image enclosures
- Configurable per-feed result limits
- Newest-first aggregation
- Custom `FeedSource` support
- Network-free unit tests for parsing and aggregation behavior
- No web-framework dependency
## Default sources
The included source set reflects the feeds used by the mature portfolio implementation:
| Category | Publication |
| --- | --- |
| Space | NASA |
| Technology | Ars Technica |
| Technology | MIT Technology Review |
| Science | Quanta Magazine |
| Science | Nautilus |
| Ideas | Aeon |
| Exploration | Atlas Obscura |
| Exploration / History | Smithsonian Magazine |
Feed availability and publisher RSS formats are external dependencies and can change independently of this project.
## Project structure
```text
rss_feed_aggregator/
├── __init__.py
├── aggregator.py
├── feeds.py
├── models.py
└── parser.py
examples/
└── basic_usage.py
tests/
├── test_aggregator.py
└── test_parser.py
```
## Installation
Create a virtual environment and install the two runtime dependencies:
```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```
## Basic usage
```python
from rss_feed_aggregator import load_articles
articles = load_articles(per_feed_limit=4)
for article in articles[:10]:
print(article.source_label)
print(article.title)
print(article.published)
print(article.link)
```
### Use your own feeds
```python
from rss_feed_aggregator import FeedSource, load_articles
feeds = [
FeedSource(
name="Example Publication",
url="https://example.com/feed.xml",
category="Research",
)
]
articles = load_articles(feeds, per_feed_limit=10)
```
Use `per_feed_limit=None` to load every entry returned by each configured feed.
## Normalized article model
Each feed entry becomes an `Article` with:
```text
title
summary
author
published
link
image
source
category
```
`source_label` is also available as a convenience property, for example `Science • Quanta Magazine`.
## Testing
The test suite does not make network requests:
```bash
python -m unittest discover -s tests -v
```
Live feed behavior is intentionally separate from the unit suite because publishers can change availability and RSS markup without warning.
## Architecture
See [`docs/architecture.md`](docs/architecture.md) for the extraction and normalization pipeline.

## Historical integration
The aggregator was developed to power a Reading interface that mixed curated external publications with first-party writing. The web interface is not part of this repository; the screenshot below is retained as context for the project that consumed the engine.

## Maintenance notes
This extraction incorporates the mature portfolio parser rather than the original single-feed prototype. In particular, it adds the broader source set, robust date fallbacks, summary cleaning, and the image extraction strategies that accumulated during real use.
Application-specific behavior from the portfolio—internal articles, featured selection, Flask/Jinja rendering, and caching—has deliberately not been copied into this package.
## License
MIT. See [`LICENSE`](LICENSE).
docs/architecture.md
# Architecture
RSS Feed Aggregator is intentionally framework-independent. It performs one job: convert multiple external RSS/Atom feeds into one normalized, newest-first collection of Python objects.
```text
Configured FeedSource objects
│
▼
feedparser
│
▼
parser.py
┌────────┼─────────┐
│ │ │
dates summaries images
└────────┼─────────┘
▼
Article
│
▼
aggregator.py
│
▼
newest-first article list
```
## Components
### `models.py`
Defines the normalized `Article` model and the `FeedSource` configuration model. Presentation layers can use `Article.source_label` when they want a combined category/publication label.
### `feeds.py`
Contains the default curated source configuration. The aggregator is not coupled to those feeds; callers may supply their own `FeedSource` iterable to `load_articles()`.
### `parser.py`
Downloads one feed and normalizes its entries. It handles:
- `published_parsed` with `updated_parsed` fallback;
- HTML removal and whitespace normalization in summaries;
- configurable summary truncation;
- Media RSS `media:content` images;
- Media RSS thumbnails;
- images embedded in content HTML;
- images embedded in summaries;
- image enclosures;
- missing titles, authors, summaries, links, dates, and images.
### `aggregator.py`
Loads each configured source, combines normalized articles, and sorts the result by publication time. The per-feed entry limit is configurable.
## Deliberately excluded
The original portfolio integration also handled internal Markdown articles, featured-content selection, request-time caching, Jinja templates, and Flask routes. That belongs to the consuming app (my portfolio), not the aggregation engine.
rss_feed_aggregator/__init__.py
from .aggregator import load_articles
from .feeds import FEEDS
from .models import Article, FeedSource
from .parser import clean_summary, extract_image, load_feed, parse_date
__all__ = [
"Article",
"FeedSource",
"FEEDS",
"clean_summary",
"extract_image",
"load_articles",
"load_feed",
"parse_date",
]
rss_feed_aggregator/models.py
from dataclasses import dataclass
from datetime import datetime
@dataclass(frozen=True, slots=True)
class FeedSource:
"""Configuration for a single RSS source."""
name: str
url: str
category: str
default_author: str | None = None
@property
def label(self) -> str:
"""Human-readable source label suitable for presentation layers."""
return f"{self.category} • {self.name}"
@dataclass(slots=True)
class Article:
"""Normalized article produced from an RSS entry."""
title: str
summary: str
author: str
published: datetime
link: str
image: str | None
source: str
category: str
@property
def source_label(self) -> str:
"""Return a display label combining category and publication."""
return f"{self.category} • {self.source}"
rss_feed_aggregator/feeds.py
from .models import FeedSource
FEEDS = [
FeedSource(
name="NASA",
url="https://www.nasa.gov/rss/dyn/breaking_news.rss",
category="Space",
default_author="NASA",
),
FeedSource(
name="Ars Technica",
url="https://feeds.arstechnica.com/arstechnica/index",
category="Technology",
default_author="Ars Technica",
),
FeedSource(
name="Quanta Magazine",
url="https://www.quantamagazine.org/feed/",
category="Science",
default_author="Quanta Magazine",
),
FeedSource(
name="Nautilus",
url="https://nautil.us/feed/",
category="Science",
default_author="Nautilus",
),
FeedSource(
name="MIT Technology Review",
url="https://www.technologyreview.com/feed/",
category="Technology",
default_author="MIT Technology Review",
),
FeedSource(
name="Smithsonian Magazine",
url="https://www.smithsonianmag.com/rss/latest_articles/",
category="Exploration / History",
default_author="Smithsonian Magazine",
),
FeedSource(
name="Aeon",
url="https://aeon.co/feed.rss",
category="Ideas",
default_author="Aeon",
),
FeedSource(
name="Atlas Obscura",
url="https://www.atlasobscura.com/feeds/latest",
category="Exploration",
default_author="Atlas Obscura",
),
]
rss_feed_aggregator/parser.py
from __future__ import annotations
from datetime import datetime
import re
from typing import Any
import feedparser
from bs4 import BeautifulSoup
from .models import Article, FeedSource
def parse_date(entry: Any) -> datetime:
"""Return the best publication timestamp available on an RSS entry.
``published_parsed`` is preferred, followed by ``updated_parsed``. Entries
without either value receive ``datetime.min`` so an undated item does not
incorrectly appear newer than dated articles when the aggregate is sorted.
"""
for field in ("published_parsed", "updated_parsed"):
parsed = _entry_get(entry, field)
if parsed:
try:
return datetime(*parsed[:6])
except (TypeError, ValueError):
continue
return datetime.min
def clean_summary(summary: str | None, max_length: int = 180) -> str:
"""Strip HTML/duplicate whitespace and truncate an RSS summary."""
if not summary:
return ""
text = BeautifulSoup(str(summary), "html.parser").get_text(" ", strip=True)
text = re.sub(r"\s+", " ", text).strip()
if len(text) <= max_length:
return text
shortened = text[:max_length].rsplit(" ", 1)[0].rstrip()
if not shortened:
shortened = text[:max_length].rstrip()
return f"{shortened}..."
def extract_image(entry: Any) -> str | None:
"""Extract an article image using common RSS/Atom conventions.
Extraction order mirrors the mature implementation used by the original
portfolio integration: Media RSS content, Media RSS thumbnail, HTML
content, HTML summary, and finally image enclosures.
"""
media_content = _entry_get(entry, "media_content") or []
if media_content:
url = _mapping_get(media_content[0], "url")
if url:
return url
media_thumbnail = _entry_get(entry, "media_thumbnail") or []
if media_thumbnail:
url = _mapping_get(media_thumbnail[0], "url")
if url:
return url
content = _entry_get(entry, "content") or []
if content:
value = _mapping_get(content[0], "value")
image = _first_html_image(value)
if image:
return image
summary = _entry_get(entry, "summary")
image = _first_html_image(summary)
if image:
return image
links = _entry_get(entry, "links") or []
for link in links:
rel = _mapping_get(link, "rel")
media_type = _mapping_get(link, "type") or ""
href = _mapping_get(link, "href")
if rel == "enclosure" and str(media_type).startswith("image/") and href:
return href
return None
def load_feed(source: FeedSource, limit: int | None = None) -> list[Article]:
"""Download one RSS feed and normalize its entries into ``Article`` objects."""
feed = feedparser.parse(source.url)
entries = list(getattr(feed, "entries", []) or [])
if limit is not None:
if limit < 0:
raise ValueError("limit must be zero or greater")
entries = entries[:limit]
articles: list[Article] = []
for entry in entries:
title = str(_entry_get(entry, "title") or "Untitled").strip()
summary = clean_summary(
_entry_get(entry, "summary") or _entry_get(entry, "description") or ""
)
author = str(
_entry_get(entry, "author")
or source.default_author
or source.name
).strip()
link = str(_entry_get(entry, "link") or "").strip()
articles.append(
Article(
title=title,
summary=summary,
author=author,
published=parse_date(entry),
link=link,
image=extract_image(entry),
source=source.name,
category=source.category,
)
)
return articles
def _entry_get(entry: Any, key: str, default: Any = None) -> Any:
if hasattr(entry, "get"):
try:
return entry.get(key, default)
except TypeError:
pass
return getattr(entry, key, default)
def _mapping_get(value: Any, key: str, default: Any = None) -> Any:
if hasattr(value, "get"):
return value.get(key, default)
return getattr(value, key, default)
def _first_html_image(html: str | None) -> str | None:
if not html:
return None
soup = BeautifulSoup(str(html), "html.parser")
image = soup.find("img")
if image:
src = image.get("src")
if src:
return str(src)
return None
rss_feed_aggregator/aggregator.py
from collections.abc import Iterable
from .feeds import FEEDS
from .models import Article, FeedSource
from .parser import load_feed
def load_articles(
feeds: Iterable[FeedSource] | None = None,
*,
per_feed_limit: int | None = 4,
) -> list[Article]:
"""Load, combine, and newest-first sort articles from configured feeds.
Args:
feeds: Optional iterable of ``FeedSource`` objects. Defaults to ``FEEDS``.
per_feed_limit: Maximum entries loaded from each feed. Use ``None`` to
load every entry returned by each source.
"""
if per_feed_limit is not None and per_feed_limit < 0:
raise ValueError("per_feed_limit must be zero or greater")
articles: list[Article] = []
for source in FEEDS if feeds is None else feeds:
articles.extend(load_feed(source, limit=per_feed_limit))
articles.sort(key=lambda article: article.published, reverse=True)
return articles
examples/basic_usage.py
from rss_feed_aggregator import load_articles
articles = load_articles(per_feed_limit=4)
for article in articles[:10]:
print(article.source_label)
print(article.title)
print(article.author)
print(article.published)
print(article.link)
print(article.image or "No feed image")
print("-" * 60)
tests/test_parser.py
from datetime import datetime
import unittest
from unittest.mock import patch
from rss_feed_aggregator.models import FeedSource
from rss_feed_aggregator.parser import (
clean_summary,
extract_image,
load_feed,
parse_date,
)
class ParserTests(unittest.TestCase):
def test_parse_date_prefers_published(self):
entry = {
"published_parsed": (2026, 8, 15, 10, 30, 0, 0, 0, 0),
"updated_parsed": (2026, 8, 16, 10, 30, 0, 0, 0, 0),
}
self.assertEqual(parse_date(entry), datetime(2026, 8, 15, 10, 30))
def test_parse_date_falls_back_to_updated(self):
entry = {
"updated_parsed": (2026, 8, 14, 8, 0, 0, 0, 0, 0),
}
self.assertEqual(parse_date(entry), datetime(2026, 8, 14, 8, 0))
def test_parse_date_uses_datetime_min_when_date_missing(self):
self.assertEqual(parse_date({}), datetime.min)
def test_clean_summary_strips_html_and_collapses_whitespace(self):
summary = "<p>Hello <strong>world</strong>.</p>\n<p>More text.</p>"
self.assertEqual(clean_summary(summary), "Hello world . More text.")
def test_clean_summary_truncates_at_word_boundary(self):
summary = "one two three four five"
self.assertEqual(clean_summary(summary, max_length=13), "one two...")
def test_extract_image_prefers_media_content(self):
entry = {
"media_content": [{"url": "https://example.test/media.jpg"}],
"media_thumbnail": [{"url": "https://example.test/thumb.jpg"}],
"summary": '<img src="https://example.test/summary.jpg">',
}
self.assertEqual(
extract_image(entry),
"https://example.test/media.jpg",
)
def test_extract_image_uses_html_summary(self):
entry = {
"summary": '<p>Story</p><img src="https://example.test/summary.jpg">',
}
self.assertEqual(
extract_image(entry),
"https://example.test/summary.jpg",
)
def test_extract_image_uses_image_enclosure(self):
entry = {
"links": [
{
"rel": "enclosure",
"type": "image/jpeg",
"href": "https://example.test/enclosure.jpg",
}
]
}
self.assertEqual(
extract_image(entry),
"https://example.test/enclosure.jpg",
)
@patch("rss_feed_aggregator.parser.feedparser.parse")
def test_load_feed_normalizes_entry(self, parse_mock):
parse_mock.return_value.entries = [
{
"title": "A story",
"summary": "<p>A useful summary.</p>",
"published_parsed": (2026, 8, 15, 9, 0, 0, 0, 0, 0),
"link": "https://example.test/story",
}
]
source = FeedSource(
name="Example",
url="https://example.test/feed.xml",
category="Science",
default_author="Example Publication",
)
articles = load_feed(source, limit=1)
self.assertEqual(len(articles), 1)
self.assertEqual(articles[0].title, "A story")
self.assertEqual(articles[0].summary, "A useful summary.")
self.assertEqual(articles[0].author, "Example Publication")
self.assertEqual(articles[0].source, "Example")
self.assertEqual(articles[0].category, "Science")
self.assertEqual(articles[0].source_label, "Science • Example")
if __name__ == "__main__":
unittest.main()
tests/test_aggregator.py
from datetime import datetime
import unittest
from unittest.mock import patch
from rss_feed_aggregator import Article, FeedSource
from rss_feed_aggregator.aggregator import load_articles
class AggregatorTests(unittest.TestCase):
@patch("rss_feed_aggregator.aggregator.load_feed")
def test_combines_and_sorts_articles_newest_first(self, load_feed_mock):
first = FeedSource("First", "https://first.test/feed", "Science")
second = FeedSource("Second", "https://second.test/feed", "Ideas")
older = Article(
title="Older",
summary="",
author="First",
published=datetime(2026, 8, 10),
link="https://first.test/older",
image=None,
source="First",
category="Science",
)
newer = Article(
title="Newer",
summary="",
author="Second",
published=datetime(2026, 8, 15),
link="https://second.test/newer",
image=None,
source="Second",
category="Ideas",
)
load_feed_mock.side_effect = [[older], [newer]]
articles = load_articles([first, second], per_feed_limit=3)
self.assertEqual([article.title for article in articles], ["Newer", "Older"])
self.assertEqual(load_feed_mock.call_count, 2)
for call in load_feed_mock.call_args_list:
self.assertEqual(call.kwargs["limit"], 3)
def test_negative_limit_is_rejected(self):
with self.assertRaises(ValueError):
load_articles([], per_feed_limit=-1)
if __name__ == "__main__":
unittest.main()
CHANGELOG.md
# Changelog
## 2026-08-15
- Expanded the original NASA-only source configuration to the curated source set used by the mature portfolio integration.
- Replaced source-specific loader duplication with the reusable `FeedSource` + `load_feed()` pipeline.
- Added publication-date fallback handling.
- Added HTML/whitespace summary normalization and configurable truncation.
- Added Media RSS, embedded HTML, summary, and enclosure image extraction strategies.
- Added fallback handling for incomplete feed entries.
- Added category metadata and `source_label` to normalized articles.
- Added configurable per-feed aggregation limits.
- Added network-free parser and aggregator unit tests.
- Kept Flask, Jinja, internal Markdown articles, featured-selection logic, and application caching outside the package.
- Added runtime dependencies, `.gitignore`, documentation refresh, and MIT license text.