Scientific Citation Retrieval
Find the papers a research paper cites, out of a corpus of 20,000. Our team of four fused eight retrievers, and about half of what we gained came from noticing that 97.6% of citations stay inside one field.
Academic
I'm Oliver, a British master's student in AI at Université Paris-Saclay. I did a year of Business Science in Cape Town, then computer science in Prague, which happened almost by chance just before ChatGPT came out, and I ended up properly hooked. Most of what I've built since is retrieval and evaluation. With a team of four I worked on predicting which papers a research paper cites, out of 20,000, and a retrieval system I built is used every day by a marketing consultancy's content team. I learned the most from checking that system's automatic quality score against blind human ratings, because the score didn't track them.
A mix of group and solo work, all of it machine learning of one kind or another, and most of it more fun than it sounds. The code for the public ones is on GitHub. Each project page carries the longer version, including what the work doesn't show.
Find the papers a research paper cites, out of a corpus of 20,000. Our team of four fused eight retrievers, and about half of what we gained came from noticing that 97.6% of citations stay inside one field.
I built the retrieval system a marketing consultancy's content team uses every day to draft posts in each client's voice. My M1 internship was the study behind it. Retrieval raised an automatic voice score in testing, but when I checked that score against blind human ratings it didn't track them, so it can't be trusted on its own. I also found my own scoring was matching posts against themselves, and corrected it in the report.
The corpus is eight clients' published material and the repository is private. The method and the numbers are described here, but the data can't be shared.
We built a VAE and DRAW from scratch to see what DRAW's steps and attention each add. DRAW builds a digit up over ten steps, and can learn to move a small attention window around as it goes. A group project of three, in PyTorch.
Reading 27 emotions out of Reddit comments, in a group of three. Fine-tuned RoBERTa reached 0.5296 macro F1. The BiLSTM we built first did no better than counting words.
Two open models were asked whether the same sentence was positive or negative, with only the politician's name swapped. The verdict changed depending on whose name it was. Asking the model politely to be fair made it worse.
Not published. The notebook needs its credentials stripped first, and the work belongs to a team of four, so releasing it is not a decision I can take on my own. Happy to walk through the method and the numbers.
Real patient data is scarce and nearly impossible to share, so the question is whether you can generate fake patients believable enough to train on. A research project of about fifteen students. I built the evaluation framework, which was merged as the project's first pull request.
The repository belongs to the supervising professor, not to me, so I can't republish it.
A web app you teach your own exercises to, by doing them in front of a webcam, and then work out against. Most of the effort went into making the teaching part feel obvious, because nobody outside machine learning has much idea what a model needs to learn from. It all runs in the browser, so no video leaves your machine. Whether the guided version actually produces better classifiers than an unguided one is a study I proposed and never ran.
Procedural terrain drawn straight into the console, as ASCII shading or coloured Unicode blocks. Three algorithms: random noise, Perlin noise, and the diamond-square fractal method. I built it about six weeks before starting the thesis, with no graphics engine, just to see whether the landscapes came out looking like landscapes.
Petr Šimůnek, Oliver Wakeford
IEEE Conference on Games 2026 (CoG), Madrid
Full paper, IEEE CoG 2026 (oral). Grew out of my BSc thesis, defended 2025. It keeps the designer's fixed heights exactly: mean-height bias about 81% lower than the blurred production baseline (0.01346 against 0.07112), at comparable runtime, and 2.1× to 23.9× faster than the harmonic baseline.
Second author. The work started as my BSc thesis, and my supervisor extended the method, ran the experiments and led the writing.
Inverse distance weighting scores 0.01351 on the same metric, a margin of about 0.4%. And total variation is worse, 933.5, against 480.9 for the harmonic baseline, so the surface is rougher.
Oliver Wakeford
Master's internship report, ACL format
Submitted August 2026. Not peer-reviewed, and not published
An internship deliverable assessed by the university, not a publication. The corpus belongs to a client and can't be shared.
Université Paris-Saclay, LISN · Paris, France
M1 finished in June 2026 with an overall mark of 15.097/20. I'm in the M2 year now: retrieval, evaluation, NLP, neuro-symbolic methods and generative models, then the internship, which is the whole second semester.
Programme syllabusSept 2025 to Aug 2027
Charles University, Prague · Prague, Czechia
Final exam graded 1 (Excellent) overall. Where I learned to code, and where I got into machine learning. Algorithms, software engineering and lots and lots of maths, with the artificial intelligence specialisation on top. My thesis, a terrain generator built for the game Engine Evolution (1.5M+ downloads), was graded 1 and extended into the IEEE paper below.
Programme, Charles UniversitySept 2022 to Sept 2025
University of Cape Town · Cape Town, South Africa
A year of business science: economics, accounting and statistics.
2021
Newest first: the M2 year I'm in now, then the M1 year behind it, then the BSc at Charles University. The M2 list is what the timetable says, so most of it is still ahead of me, labelled by teaching period.
Paris-Saclay · M2, in progress
Nine courses of three credits each, plus one soft skill, and then the internship.
Paris-Saclay · M1, complete
Mathematical Skills, Discrete Mathematics, Introduction to Algorithms, Introduction to Networking, Programming 1, Principles of Computers, Czech for Beginners I, Linear Algebra 1, Mathematical Analysis 1, Algorithms and Data Structures 1, Introduction to Linux, Linear Algebra 2, Programming 2, Extension Seminar: Algorithms and Data Structures 1, Computer Systems
Elements of AI+, Combinatorics and Graph Theory 1, Programming in Java, Algorithms and Data Structures 2, Mathematical Analysis 2, Database Systems, Propositional and Predicate Logic, Extension Seminar: Algorithms and Data Structures 2, Automata and Grammars, Probability and Statistics 1, Introduction to Artificial Intelligence, Non-procedural Programming, Ethics of AI+, Individual Software Project
Introduction to Language Technologies, Chapters from Virtual Reality, Introduction to Machine Learning with Python, Machine Learning in Computer Vision, Evolutionary Algorithms 1, Introduction to Computer Linguistics, Programming in C#, Programming and Data Processing in Python, Bachelor Thesis (Consultations), Diploma Seminar for Computer Game Development, Nature Inspired Algorithms, Natural Language Processing, Advanced Java Programming, Introduction to Computer Game Development
The thesis is in Charles University's repository, and the code for five of the eight projects above is on GitHub. If a figure here matters to something you're deciding, ask me and I'll send you the source.