Selected as 1 of 24 participants worldwide (~1.5% acceptance rate) for PAIR, a selective, fully funded 10-day program focused on AI, reasoning, and cognition for students aged 16-21.
Studied information theory, AI security (adversarial attacks and cryptographic backdoors), AI governance, neuroscience (predictive processing and consciousness), and mechanism/model design.
Worked on an AI-for-mathematical-discovery project exploring the Zarankiewicz frontier. Built a Claude Code agent-prioritisation dashboard to coordinate and rank agent tasks, as part of Creation/hackathon day: https://github.com/kseniia-strelbytska/agent-dashboard
Worked under the mentorship of the founder and Professor of Machine Learning at Cambridge University, Ferenc Huszár. Reasonable works on formal verification of software — proving it's doing exactly what’s intended.
Automated translation of system protocols from TLA+ to Verus, using LLM prover-reviewer loops gated by deterministic checks. Learned TLA+ ecosystem (weak/strong fairness, TLC, TLAPS, PlusCAl) and filtered out duplicate, vacuous, and buggy-by-design specs from the corpus. Learned formal verification and anatomy of proof in Verus. Used HDBSCAN clustering on embeddings to investigate the TLA+ dataset. Successfully translated a large corpus using parallel systemd workers on AWS EC2, submitting containerised jobs to Batch/Spot (checkpointing to S3, handling Spot reclaim, claim-and-heartbeat). Produced a per-property dataset for benchmarking and fine-tuning (RLVR) LLMs in Verus proofs — to be used by other teams in the startup.
Automated proof of system properties in Verus, to create SFT data with extracted chain-of-thought. Developed an integrity boundary between the prover and the reviewer agents using Linux UIDs, signed ledgers, and git hooks. Integrated OpenRouter model support (opencode harness), Grok (grok-build), and GLM (opencode). Red-teamed the pipeline; identified published TLA+ protocols with provably false properties, each with a counterexample; ran experiments on reviewer strength — a 48-run controlled experiment with an always-accept control gate to investigate the reviewer LLM's role in generating proofs.
Skills: Formal Verification · Verus · Temporal Logic of Actions (TLA+) · LLM Orchestration · Experimental Design · Unsupervised Learning · Rust · Python · Docker · Linux · Git · Amazon Web Services (AWS) · AWS Batch · AWS Fargate · Amazon CloudWatch · Terraform · SQLite
Skills: PyTorch · Diffusion Models · Transformer Models · Adversarial Attacks · Machine Learning