RSS Feed Aggregator

Python engine for reading RSS and Atom feeds and normalizing articles into a common newest-first model.

Python · feedparser · BeautifulSoup · RSS · Atom

README.md
# RSS Feed Aggregator

A small, framework-independent Python engine for pulling articles from curated RSS/Atom feeds, normalizing inconsistent feed metadata into one common model, and returning a combined newest-first collection.

The project began as a standalone NASA feed experiment and later evolved as a proper aggregator. This repository now contains the generic aggregation logic.

## Features

- Aggregate multiple RSS/Atom publications through one reusable pipeline
- Normalize feeds into a common `Article` dataclass
- Curated feed configuration with categories and fallback authors
- Publication-date parsing with updated-date fallback
- HTML cleanup and whitespace normalization for summaries
- Configurable summary truncation
- Multiple image extraction strategies:
  - Media RSS content
  - Media RSS thumbnails
  - embedded content images
  - summary images
  - image enclosures
- Configurable per-feed result limits
- Newest-first aggregation
- Custom `FeedSource` support
- Network-free unit tests for parsing and aggregation behavior
- No web-framework dependency

## Default sources

The included source set reflects the feeds used by the mature portfolio implementation:

| Category | Publication |
| --- | --- |
| Space | NASA |
| Technology | Ars Technica |
| Technology | MIT Technology Review |
| Science | Quanta Magazine |
| Science | Nautilus |
| Ideas | Aeon |
| Exploration | Atlas Obscura |
| Exploration / History | Smithsonian Magazine |

Feed availability and publisher RSS formats are external dependencies and can change independently of this project.

## Project structure

```text
rss_feed_aggregator/
├── __init__.py
├── aggregator.py
├── feeds.py
├── models.py
└── parser.py

examples/
└── basic_usage.py

tests/
├── test_aggregator.py
└── test_parser.py
```

## Installation

Create a virtual environment and install the two runtime dependencies:

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

## Basic usage

```python
from rss_feed_aggregator import load_articles

articles = load_articles(per_feed_limit=4)

for article in articles[:10]:
    print(article.source_label)
    print(article.title)
    print(article.published)
    print(article.link)
```

### Use your own feeds

```python
from rss_feed_aggregator import FeedSource, load_articles

feeds = [
    FeedSource(
        name="Example Publication",
        url="https://example.com/feed.xml",
        category="Research",
    )
]

articles = load_articles(feeds, per_feed_limit=10)
```

Use `per_feed_limit=None` to load every entry returned by each configured feed.

## Normalized article model

Each feed entry becomes an `Article` with:

```text
title
summary
author
published
link
image
source
category
```

`source_label` is also available as a convenience property, for example `Science • Quanta Magazine`.

## Testing

The test suite does not make network requests:

```bash
python -m unittest discover -s tests -v
```

Live feed behavior is intentionally separate from the unit suite because publishers can change availability and RSS markup without warning.

## Architecture

See [`docs/architecture.md`](docs/architecture.md) for the extraction and normalization pipeline.

![Architecture](screenshots/architecture.png)

## Historical integration

The aggregator was developed to power a Reading interface that mixed curated external publications with first-party writing. The web interface is not part of this repository; the screenshot below is retained as context for the project that consumed the engine.

![Reading Page](screenshots/reading-page.png)

## Maintenance notes

This extraction incorporates the mature portfolio parser rather than the original single-feed prototype. In particular, it adds the broader source set, robust date fallbacks, summary cleaning, and the image extraction strategies that accumulated during real use.

Application-specific behavior from the portfolio—internal articles, featured selection, Flask/Jinja rendering, and caching—has deliberately not been copied into this package.

## License

MIT. See [`LICENSE`](LICENSE).
docs/architecture.md
# Architecture

RSS Feed Aggregator is intentionally framework-independent. It performs one job: convert multiple external RSS/Atom feeds into one normalized, newest-first collection of Python objects.

```text
Configured FeedSource objects
            │
            ▼
       feedparser
            │
            ▼
        parser.py
   ┌────────┼─────────┐
   │        │         │
 dates   summaries   images
   └────────┼─────────┘
            ▼
         Article
            │
            ▼
      aggregator.py
            │
            ▼
  newest-first article list
```

## Components

### `models.py`

Defines the normalized `Article` model and the `FeedSource` configuration model. Presentation layers can use `Article.source_label` when they want a combined category/publication label.

### `feeds.py`

Contains the default curated source configuration. The aggregator is not coupled to those feeds; callers may supply their own `FeedSource` iterable to `load_articles()`.

### `parser.py`

Downloads one feed and normalizes its entries. It handles:

- `published_parsed` with `updated_parsed` fallback;
- HTML removal and whitespace normalization in summaries;
- configurable summary truncation;
- Media RSS `media:content` images;
- Media RSS thumbnails;
- images embedded in content HTML;
- images embedded in summaries;
- image enclosures;
- missing titles, authors, summaries, links, dates, and images.

### `aggregator.py`

Loads each configured source, combines normalized articles, and sorts the result by publication time. The per-feed entry limit is configurable.

## Deliberately excluded

The original portfolio integration also handled internal Markdown articles, featured-content selection, request-time caching, Jinja templates, and Flask routes. That belongs to the consuming app (my portfolio), not the aggregation engine.
rss_feed_aggregator/__init__.py
from .aggregator import load_articles
from .feeds import FEEDS
from .models import Article, FeedSource
from .parser import clean_summary, extract_image, load_feed, parse_date

__all__ = [
    "Article",
    "FeedSource",
    "FEEDS",
    "clean_summary",
    "extract_image",
    "load_articles",
    "load_feed",
    "parse_date",
]
rss_feed_aggregator/models.py
from dataclasses import dataclass
from datetime import datetime


@dataclass(frozen=True, slots=True)
class FeedSource:
    """Configuration for a single RSS source."""

    name: str
    url: str
    category: str
    default_author: str | None = None

    @property
    def label(self) -> str:
        """Human-readable source label suitable for presentation layers."""

        return f"{self.category} • {self.name}"


@dataclass(slots=True)
class Article:
    """Normalized article produced from an RSS entry."""

    title: str
    summary: str
    author: str
    published: datetime
    link: str
    image: str | None
    source: str
    category: str

    @property
    def source_label(self) -> str:
        """Return a display label combining category and publication."""

        return f"{self.category} • {self.source}"
rss_feed_aggregator/feeds.py
from .models import FeedSource


FEEDS = [
    FeedSource(
        name="NASA",
        url="https://www.nasa.gov/rss/dyn/breaking_news.rss",
        category="Space",
        default_author="NASA",
    ),
    FeedSource(
        name="Ars Technica",
        url="https://feeds.arstechnica.com/arstechnica/index",
        category="Technology",
        default_author="Ars Technica",
    ),
    FeedSource(
        name="Quanta Magazine",
        url="https://www.quantamagazine.org/feed/",
        category="Science",
        default_author="Quanta Magazine",
    ),
    FeedSource(
        name="Nautilus",
        url="https://nautil.us/feed/",
        category="Science",
        default_author="Nautilus",
    ),
    FeedSource(
        name="MIT Technology Review",
        url="https://www.technologyreview.com/feed/",
        category="Technology",
        default_author="MIT Technology Review",
    ),
    FeedSource(
        name="Smithsonian Magazine",
        url="https://www.smithsonianmag.com/rss/latest_articles/",
        category="Exploration / History",
        default_author="Smithsonian Magazine",
    ),
    FeedSource(
        name="Aeon",
        url="https://aeon.co/feed.rss",
        category="Ideas",
        default_author="Aeon",
    ),
    FeedSource(
        name="Atlas Obscura",
        url="https://www.atlasobscura.com/feeds/latest",
        category="Exploration",
        default_author="Atlas Obscura",
    ),
]
rss_feed_aggregator/parser.py
from __future__ import annotations

from datetime import datetime
import re
from typing import Any

import feedparser
from bs4 import BeautifulSoup

from .models import Article, FeedSource


def parse_date(entry: Any) -> datetime:
    """Return the best publication timestamp available on an RSS entry.

    ``published_parsed`` is preferred, followed by ``updated_parsed``. Entries
    without either value receive ``datetime.min`` so an undated item does not
    incorrectly appear newer than dated articles when the aggregate is sorted.
    """

    for field in ("published_parsed", "updated_parsed"):
        parsed = _entry_get(entry, field)
        if parsed:
            try:
                return datetime(*parsed[:6])
            except (TypeError, ValueError):
                continue

    return datetime.min


def clean_summary(summary: str | None, max_length: int = 180) -> str:
    """Strip HTML/duplicate whitespace and truncate an RSS summary."""

    if not summary:
        return ""

    text = BeautifulSoup(str(summary), "html.parser").get_text(" ", strip=True)
    text = re.sub(r"\s+", " ", text).strip()

    if len(text) <= max_length:
        return text

    shortened = text[:max_length].rsplit(" ", 1)[0].rstrip()
    if not shortened:
        shortened = text[:max_length].rstrip()

    return f"{shortened}..."


def extract_image(entry: Any) -> str | None:
    """Extract an article image using common RSS/Atom conventions.

    Extraction order mirrors the mature implementation used by the original
    portfolio integration: Media RSS content, Media RSS thumbnail, HTML
    content, HTML summary, and finally image enclosures.
    """

    media_content = _entry_get(entry, "media_content") or []
    if media_content:
        url = _mapping_get(media_content[0], "url")
        if url:
            return url

    media_thumbnail = _entry_get(entry, "media_thumbnail") or []
    if media_thumbnail:
        url = _mapping_get(media_thumbnail[0], "url")
        if url:
            return url

    content = _entry_get(entry, "content") or []
    if content:
        value = _mapping_get(content[0], "value")
        image = _first_html_image(value)
        if image:
            return image

    summary = _entry_get(entry, "summary")
    image = _first_html_image(summary)
    if image:
        return image

    links = _entry_get(entry, "links") or []
    for link in links:
        rel = _mapping_get(link, "rel")
        media_type = _mapping_get(link, "type") or ""
        href = _mapping_get(link, "href")

        if rel == "enclosure" and str(media_type).startswith("image/") and href:
            return href

    return None


def load_feed(source: FeedSource, limit: int | None = None) -> list[Article]:
    """Download one RSS feed and normalize its entries into ``Article`` objects."""

    feed = feedparser.parse(source.url)
    entries = list(getattr(feed, "entries", []) or [])

    if limit is not None:
        if limit < 0:
            raise ValueError("limit must be zero or greater")
        entries = entries[:limit]

    articles: list[Article] = []

    for entry in entries:
        title = str(_entry_get(entry, "title") or "Untitled").strip()
        summary = clean_summary(
            _entry_get(entry, "summary") or _entry_get(entry, "description") or ""
        )
        author = str(
            _entry_get(entry, "author")
            or source.default_author
            or source.name
        ).strip()
        link = str(_entry_get(entry, "link") or "").strip()

        articles.append(
            Article(
                title=title,
                summary=summary,
                author=author,
                published=parse_date(entry),
                link=link,
                image=extract_image(entry),
                source=source.name,
                category=source.category,
            )
        )

    return articles


def _entry_get(entry: Any, key: str, default: Any = None) -> Any:
    if hasattr(entry, "get"):
        try:
            return entry.get(key, default)
        except TypeError:
            pass

    return getattr(entry, key, default)


def _mapping_get(value: Any, key: str, default: Any = None) -> Any:
    if hasattr(value, "get"):
        return value.get(key, default)
    return getattr(value, key, default)


def _first_html_image(html: str | None) -> str | None:
    if not html:
        return None

    soup = BeautifulSoup(str(html), "html.parser")
    image = soup.find("img")

    if image:
        src = image.get("src")
        if src:
            return str(src)

    return None
rss_feed_aggregator/aggregator.py
from collections.abc import Iterable

from .feeds import FEEDS
from .models import Article, FeedSource
from .parser import load_feed


def load_articles(
    feeds: Iterable[FeedSource] | None = None,
    *,
    per_feed_limit: int | None = 4,
) -> list[Article]:
    """Load, combine, and newest-first sort articles from configured feeds.

    Args:
        feeds: Optional iterable of ``FeedSource`` objects. Defaults to ``FEEDS``.
        per_feed_limit: Maximum entries loaded from each feed. Use ``None`` to
            load every entry returned by each source.
    """

    if per_feed_limit is not None and per_feed_limit < 0:
        raise ValueError("per_feed_limit must be zero or greater")

    articles: list[Article] = []

    for source in FEEDS if feeds is None else feeds:
        articles.extend(load_feed(source, limit=per_feed_limit))

    articles.sort(key=lambda article: article.published, reverse=True)
    return articles
examples/basic_usage.py
from rss_feed_aggregator import load_articles


articles = load_articles(per_feed_limit=4)

for article in articles[:10]:
    print(article.source_label)
    print(article.title)
    print(article.author)
    print(article.published)
    print(article.link)
    print(article.image or "No feed image")
    print("-" * 60)
tests/test_parser.py
from datetime import datetime
import unittest
from unittest.mock import patch

from rss_feed_aggregator.models import FeedSource
from rss_feed_aggregator.parser import (
    clean_summary,
    extract_image,
    load_feed,
    parse_date,
)


class ParserTests(unittest.TestCase):
    def test_parse_date_prefers_published(self):
        entry = {
            "published_parsed": (2026, 8, 15, 10, 30, 0, 0, 0, 0),
            "updated_parsed": (2026, 8, 16, 10, 30, 0, 0, 0, 0),
        }

        self.assertEqual(parse_date(entry), datetime(2026, 8, 15, 10, 30))

    def test_parse_date_falls_back_to_updated(self):
        entry = {
            "updated_parsed": (2026, 8, 14, 8, 0, 0, 0, 0, 0),
        }

        self.assertEqual(parse_date(entry), datetime(2026, 8, 14, 8, 0))

    def test_parse_date_uses_datetime_min_when_date_missing(self):
        self.assertEqual(parse_date({}), datetime.min)

    def test_clean_summary_strips_html_and_collapses_whitespace(self):
        summary = "<p>Hello   <strong>world</strong>.</p>\n<p>More text.</p>"
        self.assertEqual(clean_summary(summary), "Hello world . More text.")

    def test_clean_summary_truncates_at_word_boundary(self):
        summary = "one two three four five"
        self.assertEqual(clean_summary(summary, max_length=13), "one two...")

    def test_extract_image_prefers_media_content(self):
        entry = {
            "media_content": [{"url": "https://example.test/media.jpg"}],
            "media_thumbnail": [{"url": "https://example.test/thumb.jpg"}],
            "summary": '<img src="https://example.test/summary.jpg">',
        }

        self.assertEqual(
            extract_image(entry),
            "https://example.test/media.jpg",
        )

    def test_extract_image_uses_html_summary(self):
        entry = {
            "summary": '<p>Story</p><img src="https://example.test/summary.jpg">',
        }

        self.assertEqual(
            extract_image(entry),
            "https://example.test/summary.jpg",
        )

    def test_extract_image_uses_image_enclosure(self):
        entry = {
            "links": [
                {
                    "rel": "enclosure",
                    "type": "image/jpeg",
                    "href": "https://example.test/enclosure.jpg",
                }
            ]
        }

        self.assertEqual(
            extract_image(entry),
            "https://example.test/enclosure.jpg",
        )

    @patch("rss_feed_aggregator.parser.feedparser.parse")
    def test_load_feed_normalizes_entry(self, parse_mock):
        parse_mock.return_value.entries = [
            {
                "title": "A story",
                "summary": "<p>A useful summary.</p>",
                "published_parsed": (2026, 8, 15, 9, 0, 0, 0, 0, 0),
                "link": "https://example.test/story",
            }
        ]

        source = FeedSource(
            name="Example",
            url="https://example.test/feed.xml",
            category="Science",
            default_author="Example Publication",
        )

        articles = load_feed(source, limit=1)

        self.assertEqual(len(articles), 1)
        self.assertEqual(articles[0].title, "A story")
        self.assertEqual(articles[0].summary, "A useful summary.")
        self.assertEqual(articles[0].author, "Example Publication")
        self.assertEqual(articles[0].source, "Example")
        self.assertEqual(articles[0].category, "Science")
        self.assertEqual(articles[0].source_label, "Science • Example")


if __name__ == "__main__":
    unittest.main()
tests/test_aggregator.py
from datetime import datetime
import unittest
from unittest.mock import patch

from rss_feed_aggregator import Article, FeedSource
from rss_feed_aggregator.aggregator import load_articles


class AggregatorTests(unittest.TestCase):
    @patch("rss_feed_aggregator.aggregator.load_feed")
    def test_combines_and_sorts_articles_newest_first(self, load_feed_mock):
        first = FeedSource("First", "https://first.test/feed", "Science")
        second = FeedSource("Second", "https://second.test/feed", "Ideas")

        older = Article(
            title="Older",
            summary="",
            author="First",
            published=datetime(2026, 8, 10),
            link="https://first.test/older",
            image=None,
            source="First",
            category="Science",
        )
        newer = Article(
            title="Newer",
            summary="",
            author="Second",
            published=datetime(2026, 8, 15),
            link="https://second.test/newer",
            image=None,
            source="Second",
            category="Ideas",
        )

        load_feed_mock.side_effect = [[older], [newer]]

        articles = load_articles([first, second], per_feed_limit=3)

        self.assertEqual([article.title for article in articles], ["Newer", "Older"])
        self.assertEqual(load_feed_mock.call_count, 2)
        for call in load_feed_mock.call_args_list:
            self.assertEqual(call.kwargs["limit"], 3)

    def test_negative_limit_is_rejected(self):
        with self.assertRaises(ValueError):
            load_articles([], per_feed_limit=-1)


if __name__ == "__main__":
    unittest.main()
requirements.txt
feedparser
beautifulsoup4
CHANGELOG.md
# Changelog

## 2026-08-15

- Expanded the original NASA-only source configuration to the curated source set used by the mature portfolio integration.
- Replaced source-specific loader duplication with the reusable `FeedSource` + `load_feed()` pipeline.
- Added publication-date fallback handling.
- Added HTML/whitespace summary normalization and configurable truncation.
- Added Media RSS, embedded HTML, summary, and enclosure image extraction strategies.
- Added fallback handling for incomplete feed entries.
- Added category metadata and `source_label` to normalized articles.
- Added configurable per-feed aggregation limits.
- Added network-free parser and aggregator unit tests.
- Kept Flask, Jinja, internal Markdown articles, featured-selection logic, and application caching outside the package.
- Added runtime dependencies, `.gitignore`, documentation refresh, and MIT license text.

Back to Home