Multilingual News Analysis Platform
- Role
- NLP / Backend Project
- Timeline
- Academic Project
- Status
- Completed
- Category
- NLP / AI
A Flask application for processing multilingual news through language detection, abstractive summarization, sentiment analysis, and content extraction from text, URLs, PDFs, and RSS feeds.
This Flask application takes news content - pasted text, a web URL, a PDF upload, or an RSS feed - and runs it through a processing pipeline: language detection first, then abstractive summarization and sentiment analysis. It supports 10+ languages and uses Hugging Face Transformers with a model fallback hierarchy to handle content that multilingual models cannot process.
The problem
News analysis tools are usually built for a single language. Processing Arabic, Hindi, Urdu, or Chinese articles through an English-only pipeline produces unusable results. This platform addresses multilingual content by detecting the language first and routing it to an appropriate Hugging Face model, with automatic fallback when a specialized model is unavailable.
Approach
Input extraction
Four input methods are supported: direct text paste, URL content extraction (via BeautifulSoup), PDF file upload (pypdf), and RSS feed parsing (feedparser). All inputs are normalised to plain text before analysis.
Language detection
langdetect identifies the language and returns a confidence score. Detection results are shown to the user before summarization begins. Languages with explicit support include English, Spanish, French, German, Italian, Portuguese, Arabic, Hindi, Urdu, and Chinese, among others.
Abstractive summarization
Hugging Face Transformers run summarization using a model hierarchy: universal model first, then a multilingual model, then a language-specific fallback. GPU is detected automatically and used when available. Summarization quality varies by language and model availability - not every language receives identical coverage.
Sentiment analysis
Sentiment is classified as positive, negative, or neutral with a confidence level (High/Medium/Low). Multilingual sentiment models are used where possible, with automatic fallback.
Application structure
The app uses Flask's factory pattern with a Config class. Structured logging and a /health endpoint are included for production readiness. The Bootstrap-based frontend provides tabbed results display and real-time validation.
Stack
Backend
NLP
Content Extraction
Frontend
What it changes
- Language detection with confidence scoring across 10+ languages.
- Abstractive multilingual summarization with automatic model fallback.
- Sentiment analysis with confidence-level classification.
- Four input methods: direct text, URL, PDF upload, RSS feed.
- GPU auto-detection and structured logging for production use.
Summarization and sentiment quality are not uniform across all languages. Model availability and accuracy vary by language. GPU availability affects processing speed significantly.