03 · Flask data application
Atelier Clothing Recommendation
A full-stack NLP application that predicts whether a customer would recommend a clothing item based on review text.
Built database-backed filters, aggregate queries, and adaptive visualisations.
How Atelier Works
Atelier is a full-stack NLP application that predicts whether a customer would recommend a clothing item based on the text of their review. It presents a women’s e-commerce catalog built from 19,662 real customer reviews, and it combines product browsing with a machine learning pipeline that classifies new reviews as Recommended or Not Recommended.
The platform brings catalog discovery, keyword search, and ML-assisted review submission into one system, so shoppers can explore items, read existing feedback, and contribute a new review without leaving the application.
1. Catalog Home and Product Discovery
The experience begins on the homepage, which acts as the central access point for the catalog. Users can browse a paginated grid of clothing items, each showing its class (for example, Dresses or Jackets), a short description, average star rating, review count, and the percentage of reviewers who would recommend it.
Department shortcuts sit at the top of the page, so users can move directly into a category or continue scrolling through the full catalog. Because there is no login requirement, anyone can explore products, search, and submit a review immediately.
This homepage gives users a clear overview of the store without needing a separate product database or shopping-cart system. Item statistics are calculated from the underlying review data, so ratings and recommendation percentages stay consistent across the site.
2. Category Browsing
Users can filter the catalog by department, including Bottoms, Dresses, Intimate, Jackets, Tops, and Trend. Selecting a department opens a category page that lists only the items in that group, along with the related clothing classes for extra context.
From here, a user can open any product to see its full details and customer reviews. This structure keeps browsing organised by clothing type, similar to a typical retail site, while still using the same review-backed product data as the rest of the application.
3. Product Detail and Review History
When a user opens an item, they see a product detail page with the clothing title, description, average rating, total number of reviews, and overall recommendation percentage. Division, department, and class tags sit alongside this summary so the item is easy to place in the catalog.
Below the product information, the page lists every customer review for that item. Each review shows its title, written feedback, star rating, a Recommended or Not Recommended badge, and the reviewer’s age group.
This page is also the starting point for the machine learning workflow. A Write a Review action takes the user into a form tied to that specific clothing item, so the new review is stored against the correct product.
4. Keyword Search Workflow
Atelier includes a catalog search bar in the site header. When a user enters a query such as “summer dress,” the Flask backend tokenises and stems the search terms using NLTK’s Porter stemmer. This means related word forms, such as “dresses” and “dress,” can still match.
Each catalog item is then scored by how many stemmed query tokens appear in its title, description, class, and department. Items with no overlap are removed, and the remaining results are ranked by match score.
This is a keyword-matching search rather than a separate recommendation engine. It helps users find relevant clothing quickly while reusing the same NLP utilities that support the review classifier.
5. Review Submission and ML Recommendation Workflow
The core of Atelier is the review workflow. When a customer wants to leave feedback on an item, they complete a form with a review title, written review, star rating, and optional age.
Once submitted, the Flask frontend sends the title and review text to the backend. The text is preprocessed — lowercased, tokenised, stripped of stopwords, and filtered — then passed into a scikit-learn pipeline. That pipeline uses bag-of-words features (CountVectorizer) and a logistic regression classifier to predict whether the reviewer would recommend the item. The model also returns a confidence score.
The user is then shown a confirmation page with the predicted label — Recommended or Not Recommended — and the model’s confidence. They can accept that label or override it if it does not match their intent. This human-in-the-loop step keeps the classifier useful as a suggestion rather than a final decision the user cannot change.
After confirmation, the review is saved to a CSV file and merged back into the live catalog data. The user receives a success page with a permanent URL for the review. From that point, the item’s review count and recommendation percentage update to include the new submission, so later visitors see the latest feedback.
A typical review follows this path:
Write review → Flask form submission → Text preprocessing → Logistic regression prediction → Confirm or override label → Save to CSV → Updated product page
6. Backend, Machine Learning, and Data Flow
Atelier uses a layered Python architecture. The Flask backend is responsible for routing, page rendering, data aggregation, and connecting the user interface to the NLP and machine learning modules.
Jinja2 templates and CSS display the catalog, forms, and prediction results. The backend loads the original Kaggle dataset (assignment3_II.csv) into pandas, groups reviews by clothing title, and calculates ratings and recommendation rates for each product. Newly submitted reviews are stored in new_reviews.csv and combined with the original dataset at runtime.
The machine learning layer lives in a dedicated module. Review title and body text are combined and preprocessed with the same function used during training, which keeps inference consistent with the model the classifier was built on. The production model is a CountVectorizer + logistic regression pipeline, serialised with joblib so the web app can load it on startup instead of retraining every time.
The production model was selected after comparing three baselines on a stratified 80/20 hold-out split:
| Approach | Accuracy | F1 |
|---|---|---|
| Count Vectorizer + Logistic Regression | 89.9% | 93.8% |
| TF-IDF + Logistic Regression | 89.7% | 93.9% |
| TF-IDF + Naive Bayes | 84.9% | 91.5% |
Count + logistic regression was chosen for production because it had the strongest accuracy while remaining easy to interpret.
A typical request follows this flow:
User action → Flask route → pandas / NLP processing → scikit-learn model or CSV storage → Jinja2 response → Updated interface
This separation between the web layer, text processing, model training, and persisted data makes the application easier to test, retrain, and extend.
7. Testing, Evaluation, and Development Workflow
The project includes unit tests for the NLP and machine learning core, written with pytest. These tests cover text preprocessing, stopword removal, stemming, search scoring, and end-to-end prediction on a small synthetic dataset. They also check that training and inference use the same preprocessing steps, which is important for a text classifier.
Model quality was evaluated separately from the web interface. A hold-out evaluation script retrains candidate pipelines, records accuracy, precision, recall, F1, and a confusion matrix, then saves the selected production model. A Jupyter notebook documents exploratory analysis, including class balance, rating distribution, text length, and side-by-side model comparison charts.
Together, these steps made it possible to justify the production model with measured results, not only with the behaviour of the website.