Skip to main content
←All work

Multilingual News Analysis Platform

Role
NLP / Backend Project
Timeline
Academic Project
Status
Completed
Category
NLP / AI

A Flask application for processing multilingual news through language detection, abstractive summarization, sentiment analysis, and content extraction from text, URLs, PDFs, and RSS feeds.

This Flask application takes news content - pasted text, a web URL, a PDF upload, or an RSS feed - and runs it through a processing pipeline: language detection first, then abstractive summarization and sentiment analysis. It supports 10+ languages and uses Hugging Face Transformers with a model fallback hierarchy to handle content that multilingual models cannot process.

01

The problem

News analysis tools are usually built for a single language. Processing Arabic, Hindi, Urdu, or Chinese articles through an English-only pipeline produces unusable results. This platform addresses multilingual content by detecting the language first and routing it to an appropriate Hugging Face model, with automatic fallback when a specialized model is unavailable.

02

Approach

  1. Input extraction

    Four input methods are supported: direct text paste, URL content extraction (via BeautifulSoup), PDF file upload (pypdf), and RSS feed parsing (feedparser). All inputs are normalised to plain text before analysis.

  2. Language detection

    langdetect identifies the language and returns a confidence score. Detection results are shown to the user before summarization begins. Languages with explicit support include English, Spanish, French, German, Italian, Portuguese, Arabic, Hindi, Urdu, and Chinese, among others.

  3. Abstractive summarization

    Hugging Face Transformers run summarization using a model hierarchy: universal model first, then a multilingual model, then a language-specific fallback. GPU is detected automatically and used when available. Summarization quality varies by language and model availability - not every language receives identical coverage.

  4. Sentiment analysis

    Sentiment is classified as positive, negative, or neutral with a confidence level (High/Medium/Low). Multilingual sentiment models are used where possible, with automatic fallback.

  5. Application structure

    The app uses Flask's factory pattern with a Config class. Structured logging and a /health endpoint are included for production readiness. The Bootstrap-based frontend provides tabbed results display and real-time validation.

03

Stack

Backend

PythonFlask 2.3.3

NLP

Hugging Face Transformers 4.33.2langdetect

Content Extraction

BeautifulSoup4pypdffeedparser

Frontend

BootstrapJavaScript
04

What it changes

  • Language detection with confidence scoring across 10+ languages.
  • Abstractive multilingual summarization with automatic model fallback.
  • Sentiment analysis with confidence-level classification.
  • Four input methods: direct text, URL, PDF upload, RSS feed.
  • GPU auto-detection and structured logging for production use.

Summarization and sentiment quality are not uniform across all languages. Model availability and accuracy vary by language. GPU availability affects processing speed significantly.

05

About the author

I'm Aaqib Shaikh, a software engineer based in Karachi, Pakistan. I build full-stack web applications, AI systems, and machine-learning tools, from architecture to deployment, and studied Computer Science at Iqra University. The resume has the full background, more work is on the projects page, and the contact page is the way to reach me. Code lives on GitHub.