Python Projects for Final-Year Students: 10 Ideas with Scope, Stack and a 4-Week Build Plan
A final-year project has to do two jobs at once. It has to pass a viva, where an external examiner asks why you chose this and what the alternatives were. And it has to survive a placement interview, where an engineer asks what breaks when the data doubles. Most final-year projects fail the second job because they were scoped for the first: a big title, a thin implementation, no tests, never deployed.
This page has ten project ideas scoped the other way round. Each one is small enough to finish, has a real dataset or a real user, and has one part that is genuinely hard. For each: the scope (what is in, what is out), the stack, using only what PyRun teaches (Python standard library, pandas, Flask, FastAPI, SQLite, scikit-learn for the two ML ideas), what makes it final-year worthy rather than a weekend script, and the viva question you should expect. The last section is a four-week build plan for one of them, day by day.
Every code snippet was run on Python 3.12 before publishing. Flask, FastAPI and scikit-learn run on your own machine, not in the browser editor; pandas and the standard library run in both.
Before you pick. If you have not built anything yet, start with the placement projects hub: ten smaller projects from easy to impressive. The ideas below assume you have finished two or three of those. For the Python itself, the 10-week placement roadmap puts the lessons in order.
What makes a project "final-year"
Four things, and every idea below has all four:
- A real data source or real users. Your college's timetable, your own bank statement, a public dataset, your hostel's mess. Not
sample_data.csvwith ten rows you typed. - One hard part. Authentication, a scheduling constraint, an evaluation metric, a graph algorithm. The rest can be plain CRUD.
- Tests and a way to run it in one command.
pytestand aREADMEwithpip install -r requirements.txtandpython main.py. - A measurement. Requests per second, accuracy on held-out data, seconds saved per week. One honest number you can defend.
The 10 ideas
1. Attendance and internal-marks portal
Scope, in: faculty log in, mark attendance per lecture, enter internal marks; students log in and see only their own record; a CSV export per subject. Out: timetable generation, SMS alerts, mobile app.
Stack: Flask, SQLite, Jinja templates, werkzeug.security for password hashing, pytest.
Final-year worthy because: two roles with different permissions is the hard part. Getting "a student cannot see another student's marks" right, and proving it with a test, is exactly what backend interviewers ask about.
Lessons: Your first Flask web server, SQLite with sqlite3, SQL joins, Auth: password hashing and JWTs.
Viva question: "How do you stop a student editing the URL to another roll number?" Check the session's user against the requested record on the server, every time; never trust the URL.
2. Placement cell job board API
Scope, in: companies, drives, eligibility rules (CGPA, branch, backlog count), student applications; a REST API with token auth; a daily eligibility report as CSV. Out: a polished frontend (FastAPI's /docs page is your demo), email notifications.
Stack: FastAPI, pydantic, SQLite via SQLAlchemy, pandas for the report, pytest with TestClient.
Final-year worthy because: eligibility rules are real business logic with edge cases (CGPA 6.99 vs 7.0, "no active backlog" vs "no backlog ever"). Writing them as pure functions and testing them is the hard part. This is the project the four-week plan below builds.
Lessons: REST APIs with FastAPI, SQLAlchemy ORM, Database migrations with Alembic, Testing with pytest.
from dataclasses import dataclass
@dataclass(frozen=True)
class Drive:
company: str
min_cgpa: float
branches: frozenset[str]
max_backlogs: int = 0
@dataclass(frozen=True)
class Student:
roll: str
branch: str
cgpa: float
backlogs: int
def is_eligible(s: Student, d: Drive) -> tuple[bool, str]:
if s.branch not in d.branches:
return False, "branch not eligible"
if s.cgpa < d.min_cgpa:
return False, f"cgpa {s.cgpa} below {d.min_cgpa}"
if s.backlogs > d.max_backlogs:
return False, f"{s.backlogs} backlog(s), max {d.max_backlogs}"
return True, "eligible"
drive = Drive("Acme Systems", 7.0, frozenset({"CSE", "IT", "ECE"}), max_backlogs=0)
for st in [Student("21CS001", "CSE", 7.0, 0), Student("21CS002", "CSE", 6.99, 0),
Student("21ME003", "ME", 8.5, 0), Student("21IT004", "IT", 7.8, 1)]:
print(st.roll, is_eligible(st, drive))
Output:
21CS001 (True, 'eligible')
21CS002 (False, 'cgpa 6.99 below 7.0')
21ME003 (False, 'branch not eligible')
21IT004 (False, '1 backlog(s), max 0')
Returning the reason, not just False, is what makes the report useful and the viva easy.
Viva question: "Why FastAPI over Flask?" Automatic validation from type hints and a generated /docs page; Flask is fine too, and you should be able to say what you would lose.
3. Mess and canteen ordering with inventory
Scope, in: a menu with stock counts, orders that decrement stock atomically, a daily consumption report, a low-stock alert list. Out: payments, delivery tracking.
Stack: FastAPI or Flask, SQLite with transactions, pandas for the report.
Final-year worthy because: two people ordering the last plate at the same time is a concurrency problem. Solving it with a transaction and a WHERE stock > 0 update, and writing the test that proves it, is a story most freshers cannot tell.
Lessons: SQLite with sqlite3, Concurrency: threads vs asyncio, groupby and aggregation.
import sqlite3
conn = sqlite3.connect(":memory:")
conn.execute("CREATE TABLE menu (item TEXT PRIMARY KEY, stock INTEGER NOT NULL)")
conn.execute("INSERT INTO menu VALUES ('thali', 1)")
def order(item):
with conn: # one transaction
cur = conn.execute(
"UPDATE menu SET stock = stock - 1 WHERE item = ? AND stock > 0", (item,))
return cur.rowcount == 1 # 1 row changed = we got it
print(order("thali"), order("thali")) # True False
print(conn.execute("SELECT stock FROM menu").fetchone()[0]) # 0
The second order fails because the WHERE stock > 0 guard and the update happen in one statement; there is no window where two orders both see stock = 1.
Viva question: "What if you read the stock, checked it in Python, then wrote it back?" Race condition. Explain why the single UPDATE ... WHERE is safe.
4. UPI and bank statement analyser
Scope, in: import CSV statements from two or more banks with different column names, normalise them, categorise merchants with rules, show monthly trends and a "where did the money go" breakdown. Out: bank API integration, predictions.
Stack: pandas, matplotlib, Flask for a small dashboard, regex for merchant names.
Final-year worthy because: the hard part is data cleaning. Two banks, two date formats, one with debit and credit in separate columns and one with a signed amount. Handling that generally, with tests, is what data roles hire for.
Lessons: Cleaning messy data, Combining tables: merge, join and concat, Time series with pandas, Plotting with matplotlib and seaborn, Regex data extraction.
import io
import pandas as pd
bank_a = pd.read_csv(io.StringIO("""Date,Narration,Debit,Credit
02/09/2026,UPI-SWIGGY-BLR,240,
03/09/2026,SALARY SEP,,42000
"""))
bank_b = pd.read_csv(io.StringIO("""txn_date,description,amount
2026-09-05,UPI/ZOMATO/ORDER,-310
2026-09-06,UPI/IRCTC/TICKET,-1450
"""))
a = pd.DataFrame({
"date": pd.to_datetime(bank_a["Date"], format="%d/%m/%Y"),
"desc": bank_a["Narration"],
"amount": bank_a["Credit"].fillna(0) - bank_a["Debit"].fillna(0),
})
b = pd.DataFrame({
"date": pd.to_datetime(bank_b["txn_date"], format="%Y-%m-%d"),
"desc": bank_b["description"],
"amount": bank_b["amount"],
})
txns = pd.concat([a, b], ignore_index=True).sort_values("date")
rules = {"food": r"SWIGGY|ZOMATO", "travel": r"IRCTC|OLA|UBER", "income": r"SALARY"}
txns["category"] = "other"
for cat, pattern in rules.items():
txns.loc[txns["desc"].str.contains(pattern, case=False, regex=True), "category"] = cat
print(txns[["date", "amount", "category"]].to_string(index=False))
print(txns[txns["amount"] < 0].groupby("category")["amount"].sum())
Output:
date amount category
2026-09-02 -240.0 food
2026-09-03 42000.0 income
2026-09-05 -310.0 food
2026-09-06 -1450.0 travel
category
food -550.0
travel -1450.0
Name: amount, dtype: float64
Viva question: "How do you handle a merchant your rules do not know?" It lands in other; show the top unknown merchants each month so the rules grow. Never silently drop rows.
5. Personal study planner with spaced repetition
Scope, in: subjects and topics, a daily review queue, the SM-2 scheduling algorithm (the one Anki uses), a streak and progress view. Out: flashcard images, sync across devices.
Stack: Python standard library (sqlite3, datetime, dataclasses), Flask or a CLI.
Final-year worthy because: you implement a published algorithm and can explain its parameters. That is rarer than a CRUD app and interviewers get curious.
Lessons: Dates and Time, Dataclasses, SQLite with sqlite3.
from dataclasses import dataclass
from datetime import date, timedelta
@dataclass
class Card:
topic: str
interval: int = 1 # days until next review
ease: float = 2.5
reps: int = 0
due: date = date(2026, 10, 1)
def review(card: Card, quality: int, today: date) -> Card:
"""SM-2: quality 0-5, below 3 means forgot."""
if quality < 3:
card.reps, card.interval = 0, 1
else:
card.interval = 1 if card.reps == 0 else 6 if card.reps == 1 else round(card.interval * card.ease)
card.reps += 1
card.ease = max(1.3, card.ease + 0.1 - (5 - quality) * (0.08 + (5 - quality) * 0.02))
card.due = today + timedelta(days=card.interval)
return card
c = Card("SQL joins")
for day_offset, q in [(0, 4), (1, 5), (7, 3)]:
review(c, q, date(2026, 10, 1) + timedelta(days=day_offset))
print(f"interval {c.interval:>2}d ease {c.ease:.2f} due {c.due:%d/%m/%Y}")
Output:
interval 1d ease 2.50 due 02/10/2026
interval 6d ease 2.60 due 08/10/2026
interval 16d ease 2.46 due 24/10/2026
Viva question: "Why does the interval grow multiplicatively?" Because recall probability decays roughly exponentially, so each successful review earns a proportionally longer gap.
6. Library management with fines and reservations
Scope, in: books with copies, members, issue and return, fine calculation with a grace period, a reservation queue for books with no copies free, overdue report. Out: barcode scanning, payment collection.
Stack: Python OOP with dataclasses and custom exceptions, SQLite, Flask for the desk UI, pytest.
Final-year worthy because: the reservation queue is a real data-structure choice (deque), the fine rules have edge cases (holidays, grace days, a cap), and the whole thing is testable without a browser.
Lessons: Classes and Objects, Inheritance and super(), Queues, Deques and BFS foundations, Errors and try/except.
Viva question: "Why not a list for the reservation queue?" list.pop(0) is O(n); deque.popleft() is O(1). Small here, but the examiner is checking whether you know.
7. Log monitoring and alerting tool
Scope, in: tail one or more log files, parse them with regex, count errors per minute, alert (print, write to a file, or send a webhook) when a threshold is crossed, a summary report on exit. Out: a web dashboard, distributed collection.
Stack: Python standard library only: re, collections, argparse, logging, pathlib, generators.
Final-year worthy because: it processes an unbounded stream with constant memory, which is the same problem real monitoring tools solve. No framework to hide behind; the code is the project.
Lessons: Regular Expressions, Generators, Building CLI tools with argparse, Logging in production, Performance profiling.
from collections import deque
from datetime import datetime, timedelta
def alert_on_burst(events, threshold=3, window=timedelta(seconds=60)):
"""Yield the timestamp at which the count of events in the trailing window hits threshold."""
recent = deque()
for ts in events:
recent.append(ts)
while recent and ts - recent[0] > window:
recent.popleft()
if len(recent) == threshold:
yield ts
base = datetime(2026, 9, 24, 10, 0, 0)
errors = [base + timedelta(seconds=s) for s in (0, 20, 45, 200, 210, 215)]
for t in alert_on_burst(errors):
print("ALERT 3 errors within 60s at", t.strftime("%d/%m/%Y %H:%M:%S"))
Output:
ALERT 3 errors within 60s at 24/09/2026 10:00:45
ALERT 3 errors within 60s at 24/09/2026 10:03:35
A sliding window over time, with a deque, in constant memory. The sliding window lesson is the same idea on arrays.
Viva question: "What is the memory usage after a million lines?" Bounded by the window: only events inside the last 60 seconds are kept.
8. Lecture-notes search engine
Scope, in: index a folder of text or markdown notes, split them into chunks, rank chunks for a query with TF-IDF and cosine similarity, show the top five with the file and line. Out: PDF parsing (add it in week 4 if time allows), a large language model.
Stack: Python standard library and numpy. Optionally the PyRun GenAI lessons for embeddings, but the TF-IDF version is fully explainable in a viva, which is the point.
Final-year worthy because: you build a retrieval system from first principles and can explain every number it produces. This is the honest version of the "AI project" everyone else is claiming.
Lessons: Chunking documents for retrieval, Building a tiny vector store, Text embeddings and cosine similarity, NumPy: arrays at speed.
import math
import re
from collections import Counter
docs = {
"dbms.md": "a transaction is atomic consistent isolated durable; commit or rollback",
"os.md": "a process has its own memory; a thread shares memory with its process",
"python.md": "a generator yields values lazily; a list holds all values in memory",
}
def tokens(text):
return re.findall(r"[a-z]+", text.lower())
tf = {name: Counter(tokens(text)) for name, text in docs.items()}
df = Counter(word for c in tf.values() for word in c)
N = len(docs)
idf = {w: math.log(N / df[w]) + 1 for w in df} # +1 so shared words still count a little
def vector(counter):
return {w: n * idf.get(w, 1.0) for w, n in counter.items()}
def cosine(a, b):
dot = sum(a[w] * b[w] for w in a.keys() & b.keys())
na, nb = math.sqrt(sum(v * v for v in a.values())), math.sqrt(sum(v * v for v in b.values()))
return dot / (na * nb) if na and nb else 0.0
def search(query, k=2):
q = vector(Counter(tokens(query)))
scored = sorted(((cosine(q, vector(c)), name) for name, c in tf.items()), reverse=True)
return [(name, round(score, 3)) for score, name in scored[:k]]
print(search("thread memory"))
print(search("generator vs list"))
Output:
[('os.md', 0.398), ('python.md', 0.106)]
[('python.md', 0.381), ('os.md', 0.0)]
Viva question: "Why does memory match both os.md and python.md?" Because it appears in both, so its IDF is low; thread appears only in os.md, so it dominates. Show the numbers.
9. Exam result predictor with honest evaluation
Scope, in: a dataset of past students (attendance, internal marks, assignment scores, final result), a model that predicts pass or fail and a second one that predicts the final mark, a proper train/test split, a report of precision, recall and the baseline it beats. Out: a live prediction portal, deep learning.
Stack: pandas, scikit-learn, matplotlib.
Final-year worthy because: the hard part is evaluation, not modelling. Most student ML projects report 97% accuracy on the training set. Yours reports held-out metrics against a baseline, and explains what the model would get wrong.
Lessons: Your first machine-learning model, Train/test split, Classification, Overfitting and honest evaluation, Decision trees and feature importance.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, recall_score
rng = np.random.default_rng(42)
n = 400
attendance = rng.uniform(40, 100, n)
internal = rng.uniform(5, 30, n)
# pass if a weighted score clears a bar, with some noise
score = 0.5 * attendance + 2.0 * internal + rng.normal(0, 8, n)
passed = (score > 75).astype(int)
X = np.column_stack([attendance, internal])
X_train, X_test, y_train, y_test = train_test_split(X, passed, test_size=0.25, random_state=0)
model = LogisticRegression().fit(X_train, y_train)
pred = model.predict(X_test)
baseline = max(y_test.mean(), 1 - y_test.mean()) # always predict the majority class
print(f"baseline (majority class): {baseline:.2f}")
print(f"model accuracy on held-out: {accuracy_score(y_test, pred):.2f}")
print(f"recall for 'fail' students: {recall_score(y_test, pred, pos_label=0):.2f}")
Output (scikit-learn 1.9, synthetic data with a fixed seed):
baseline (majority class): 0.61
model accuracy on held-out: 0.91
recall for 'fail' students: 0.92
The synthetic data is only there to make the snippet runnable; your project uses real (anonymised) records. The point is the three lines of output: baseline, held-out accuracy, and recall on the class that matters (students at risk).
Viva question: "Your accuracy is 91%. Is that good?" Only relative to the 61% baseline, and only if recall on the failing class is acceptable, because those are the students the model exists to find.
10. City bus route planner
Scope, in: stops and routes for one city or one campus shuttle network as a graph, shortest path by number of stops (BFS) and by travel time (Dijkstra), a FastAPI endpoint that returns the route, tests on a small hand-checked network. Out: live bus positions, maps UI.
Stack: Python standard library (collections.deque, heapq), FastAPI.
Final-year worthy because: it is a graph problem on real data (most city transport corporations publish route lists), with a clear correctness test and a clear performance question.
Lessons: Graph representations, Graph BFS, Heaps and priority queues, REST APIs with FastAPI.
import heapq
from collections import deque
# stop -> [(neighbour, minutes)]
network = {
"Station": [("Market", 6), ("College", 15)],
"Market": [("Station", 6), ("Hospital", 5), ("College", 4)],
"Hospital": [("Market", 5), ("College", 3)],
"College": [("Station", 15), ("Market", 4), ("Hospital", 3)],
}
def fewest_stops(src, dst):
prev, queue = {src: None}, deque([src])
while queue:
node = queue.popleft()
if node == dst:
break
for nxt, _ in network[node]:
if nxt not in prev:
prev[nxt] = node
queue.append(nxt)
path, node = [], dst
while node is not None:
path.append(node); node = prev[node]
return path[::-1]
def fastest(src, dst):
best = {src: 0}
heap = [(0, src, [src])]
while heap:
t, node, path = heapq.heappop(heap)
if node == dst:
return t, path
for nxt, mins in network[node]:
if t + mins < best.get(nxt, float("inf")):
best[nxt] = t + mins
heapq.heappush(heap, (t + mins, nxt, path + [nxt]))
print(fewest_stops("Station", "Hospital"))
print(fastest("Station", "College"))
Output:
['Station', 'Market', 'Hospital']
(10, ['Station', 'Market', 'College'])
The direct Station to College bus takes 15 minutes; changing at Market takes 10. That kind of result is what makes the demo land.
Viva question: "Why does BFS give the fewest stops but not the fastest route?" BFS treats every edge as equal; Dijkstra orders by accumulated cost. Know when each applies.
A 4-week build plan: the placement cell job board API (#2)
Assume about two hours on weekdays and five on weekends, roughly 20 hours a week. The plan is for one person; a team of two splits the API and the report.
Week 1: data model and plain CRUD
- Day 1 to 2: write the eligibility rules as pure functions with the dataclasses above. Write ten tests for them first: the 6.99 CGPA case, the wrong branch, one backlog, the boundary values. This is the part the viva will ask about, so it gets tested before anything else exists.
- Day 3 to 4: SQLAlchemy models for
Student,Company,Drive,Application. Acreate_db.pythat builds a SQLite file and loads 50 students from a CSV. The SQLAlchemy ORM lesson covers the model syntax. - Day 5 to 7: FastAPI routes: list drives, get a drive, create a drive, list students. No auth yet. Open
/docsand click through every route. Commit at the end of every day.
Exit check: pytest passes, uvicorn main:app starts, and /docs shows six routes.
Week 2: eligibility, applications and auth
- Day 8 to 9:
GET /drives/{id}/eligiblereturns the eligible students with the reason for every ineligible one. Reuse the week-1 functions; do not duplicate the rules in SQL. - Day 10 to 11:
POST /drives/{id}/applyfor a student. Reject if ineligible, reject duplicates, return 201 on success. Write the tests for the rejections. - Day 12 to 14: token auth with two roles: placement officer (can create drives) and student (can apply, sees only their own applications). Hash passwords with bcrypt; sign the token as the auth lesson shows. Write the test that proves a student cannot read another student's applications.
Exit check: an unauthenticated request gets 401, a student creating a drive gets 403, and every rejection has a test.
Week 3: the report, migrations and deployment
- Day 15 to 16: the daily report with pandas: applications per drive, eligible-but-not-applied students per branch, exported as CSV and as a summary table on
GET /reports/daily. The groupby lesson has the exact operations. - Day 17 to 18: add Alembic so a schema change (say, a
resume_urlcolumn) is a migration, not a delete-the-database. The Alembic lesson and the migrations post coverrevision --autogenerateandupgrade head. - Day 19 to 21: deploy. The deploying a Python web app lesson and the free hosting comparison cover the options; a free tier with SQLite on disk is enough for a demo. Measure one number: requests per second on the eligibility endpoint with 50 students, using a simple loop and
time.perf_counter(). Write it down with the date.
Exit check: a public URL whose /docs page works on your phone, and one measured number.
Week 4: documentation, hardening and the viva
- Day 22 to 23: README: one paragraph, a screenshot of
/docs, the three commands to run it, the measured number, and a "what I would do next" list. Addrequirements.txtand a.gitignorethat excludes the database file and.env. - Day 24 to 25: hardening pass. What happens with an empty CSV? A CGPA of 11? A drive with no branches? Add the tests, fix what breaks. Run everything once on a fresh clone in a new virtual environment; half of all demo failures are "works on my machine".
- Day 26 to 27: viva preparation. Write one-paragraph answers to: why FastAPI, why SQLite, how eligibility is tested, how you stop cross-student access, what breaks at 10,000 students, and what you would change. Rehearse the ten-minute pitch from the presentation section of the projects hub.
- Day 28: rest, then a final read of the README as if you were the examiner.
Exit check: a stranger can clone the repo, run three commands, and see it work.
Choosing between them
- Backend or software roles: #1, #2, #3, #6.
- Data or analyst roles: #4, #9, then #8.
- You want the algorithm to be the story: #5, #7, #8, #10.
- Solo and short on time: #7 or #5. Standard library only, nothing to deploy, still genuinely hard.
Whichever you pick, finish it. An examiner and an interviewer both prefer a small, complete, tested project with one honest measurement to a large one that "mostly works".
Also on PyRun: the placement projects hub for smaller warm-up projects, the interview practice page for the coding round, the cheat sheets for the syntax you keep forgetting, and the level test to find the lesson to start at. The first five lessons are free, no account or card needed; the rest of the curriculum, including the Flask, FastAPI, SQL, pandas and ML lessons linked above, is part of the paid plan with a 7-day trial.