<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
<title>far.in.net</title>
<subtitle>Writing by Matthew Farrugia-Roberts</subtitle>
<link href="https://far.in.net/feed.xml" rel="self"/>
<link href="https://far.in.net/"/>
<id>https://far.in.net/feed.xml</id>
<updated>2026-07-26T00:00:00Z</updated>
<author><name>Matthew Farrugia-Roberts</name></author>
<entry>
<title>Time to write</title>
<link href="https://far.in.net/time-to-write"/>
<id>https://far.in.net/time-to-write</id>
<updated>2026-07-26T00:00:00Z</updated>
<summary>I’m writing from the liminal space between London and San Francisco. Time works differently here (roughly speaking, three hours become ten). There’s also no internet, and that means no agents or students to prompt. Seems like the perfect opportunity to prime myself for my visit to Berkeley. Let’s see what I can make of it.</summary>
</entry>
<entry>
<title>Understanding epiplexity</title>
<link href="https://far.in.net/epiplexity"/>
<id>https://far.in.net/epiplexity</id>
<updated>2026-07-08T00:00:00Z</updated>
<summary>Everyone seems to be talking about the paper by Finzi et al. (2026), “From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence.” I finally got around to studying it this week.</summary>
</entry>
<entry>
<title>Variational inference</title>
<link href="https://far.in.net/variational-inference"/>
<id>https://far.in.net/variational-inference</id>
<updated>2026-06-30T00:00:00Z</updated>
<summary>In this short note, I recall variational inference and the evidence lower bound (ELBO). In contrast to the introductions I have seen, I start by describing variational inference as a loss minimisation problem in its own right without reference to the ELBO, and then derive the ELBO afterwards.</summary>
</entry>
<entry>
<title>Subscribing to arXiv</title>
<link href="https://far.in.net/subscribing-to-arxiv"/>
<id>https://far.in.net/subscribing-to-arxiv</id>
<updated>2026-06-29T00:00:00Z</updated>
<summary>How do you find out about new papers? Do you curate an academic Twitter feed? Do you set up alerts on Google Scholar for specific keywords? Do you keep an eye on the #papers channel in each of the Slacks/Discords you frequent? Do you ask an LLM for literature recommendations?</summary>
</entry>
<entry>
<title>What is Life? by Erwin Schrödinger</title>
<link href="https://far.in.net/what-is-life"/>
<id>https://far.in.net/what-is-life</id>
<updated>2026-03-20T00:00:00Z</updated>
<summary>This post spun out from my reading log.</summary>
</entry>
<entry>
<title>What is complexity?</title>
<link href="https://far.in.net/complexity"/>
<id>https://far.in.net/complexity</id>
<updated>2026-03-16T00:00:00Z</updated>
<summary>What is complexity? Etymologically, the prefix “com-” traces back to Latin, meaning “with, together,” and “plex” traces back to PIE, meaning “to plait,” as in, “to entwine.” Complex entities are those that are formed by entwining multiple elements together.</summary>
</entry>
<entry>
<title>On how to choose what to work on</title>
<link href="https://far.in.net/quotes-on-what-to-work-on"/>
<id>https://far.in.net/quotes-on-what-to-work-on</id>
<updated>2026-03-16T00:00:00Z</updated>
<summary>A poem “The Church-Porch,” by George Herbert (1593–1633) (link).</summary>
</entry>
<entry>
<title>Replacing Guilt by Nate Soares</title>
<link href="https://far.in.net/replacing-guilt"/>
<id>https://far.in.net/replacing-guilt</id>
<updated>2026-03-07T00:00:00Z</updated>
<summary></summary>
</entry>
<entry>
<title>Transformers: Perceptrons in disguise</title>
<link href="https://far.in.net/perceptrons-in-disguise"/>
<id>https://far.in.net/perceptrons-in-disguise</id>
<updated>2026-03-05T00:00:00Z</updated>
<summary>Attention is a key element of the transformer architecture. It can be understood from several related perspectives: parallel soft search, reading from / writing to residual streams, loosely coupled QK/OV circuits, and message passing.</summary>
</entry>
<entry>
<title>Switching to JAX</title>
<link href="https://far.in.net/jax"/>
<id>https://far.in.net/jax</id>
<updated>2026-03-01T00:00:00Z</updated>
<summary>JAX is a Python library for writing and transforming array programs, such as arise in the course of deep learning research. The JAX API for writing array programs is similar to that of NumPy, but the JAX API for transforming array programs is utterly unique, effortlessly powerful, and delightfully elegant.</summary>
</entry>
<entry>
<title>The tempered posterior</title>
<link href="https://far.in.net/tempered-posterior"/>
<id>https://far.in.net/tempered-posterior</id>
<updated>2026-02-27T00:00:00Z</updated>
<summary>Most of the results in singular learning theory are developed in the context of a generalised form of Bayesian inference, in which the posterior distribution has an additional (inverse) temperature parameter. In this note, I introduce this so-called “tempered posterior” and contrast it with the usual Bayesian posterior. I also discuss various motivations for departing from the Bayesian suggestion, including a derivation of the tempered posterior using the principle of maximum entropy.</summary>
</entry>
<entry>
<title>Generalised readings</title>
<link href="https://far.in.net/readings-2019-to-2025"/>
<id>https://far.in.net/readings-2019-to-2025</id>
<updated>2026-02-01T00:00:00Z</updated>
<summary>Once upon a time, at an end-of-year teaching team dinner, I was talking with a friend of mine about New Year’s Resolutions. Every year, she said, she set herself the goal of reading 52 books the following year. What an amazing achievement! To read one book a week! I realised I wanted to spend substantially more time reading and, one day, achieve this myself.</summary>
</entry>
<entry>
<title>One size fits all</title>
<link href="https://far.in.net/one-size-fits-all"/>
<id>https://far.in.net/one-size-fits-all</id>
<updated>2026-01-02T00:00:00Z</updated>
<summary>Adrianne ran her fingers over the darts and let out a sigh of relief. The dress was a perfect fit! Part of her wasn’t surprised, of course: she had followed the pattern to the letter and measured every cut three times! But the rest of her wasn’t ready to believe it without proof.</summary>
</entry>
<entry>
<title>Freeing 118 GB of W&amp;B experiment data from a broken binary format</title>
<link href="https://far.in.net/free-wandb"/>
<id>https://far.in.net/free-wandb</id>
<updated>2025-12-29T00:00:00Z</updated>
<summary>I spent the final weeks of 2024 migrating 5,203 deep learning experiment logs out of the Weights &amp; Biases (W&amp;B) cloud and the undocumented .wandb binary file format so that I could more flexibly and efficiently analyse the data in the course of writing a paper.</summary>
</entry>
<entry>
<title>Schwartz distributions</title>
<link href="https://far.in.net/schwartz-distributions"/>
<id>https://far.in.net/schwartz-distributions</id>
<updated>2025-08-17T00:00:00Z</updated>
<summary>This is a technical note intended as an elementary introduction to Schwartz distributions and their calculus. Schwartz distributions are a concept from functional analysis used in singular learning theory to define the density of states, a key step in the bridge between algebraic geometry and Bayesian statistics.</summary>
</entry>
<entry>
<title>Blowing up</title>
<link href="https://far.in.net/blowing-up"/>
<id>https://far.in.net/blowing-up</id>
<updated>2025-08-10T00:00:00Z</updated>
<summary>Blowing up is a concept from algebraic geometry. Blowing up turns a neighbourhood in a Euclidean space into a more complex manifold (called the ‘blow up’ of the original neighbourhood). The transformation preserves the Euclidean geometry in most parts of the neighbourhood, but expands a particular subspace to have extra dimensions. This can help to resolve singularities, with applications in singular learning theory.</summary>
</entry>
<entry>
<title>Instrumental/intrinsic value ambiguity</title>
<link href="https://far.in.net/value-ambiguity"/>
<id>https://far.in.net/value-ambiguity</id>
<updated>2025-07-25T00:00:00Z</updated>
<summary>In this note, I introduce a toy environment that demonstrates instrumental/intrinsic value ambiguity. An agent repeatedly faces choices between two kinds of actions: intrinsically valuable actions (which are directly rewarded), and instrumentally valuable actions (which are not directly rewarded, but enable later intrinsically valuable actions). Crucially, I show how training in a subset of environment configurations creates a situation where the agent can’t reliably distinguish between these two kinds of value.</summary>
</entry>
<entry>
<title>Smile!</title>
<link href="https://far.in.net/smile"/>
<id>https://far.in.net/smile</id>
<updated>2025-07-15T00:00:00Z</updated>
<summary>I’m not sure if it was just part of being an introvert or a product of a childhood spent playing Pokémon games, but for a large part of my early adulthood, I was slightly scared of locking eyes with strangers in public.</summary>
</entry>
<entry>
<title>Turing trees</title>
<link href="https://far.in.net/turing-trees"/>
<id>https://far.in.net/turing-trees</id>
<updated>2025-06-15T00:00:00Z</updated>
<summary>I sometimes think about models of computation as an infinite network of infinite binary trees. I call the trees Turing trees and the network the Turing web. These objects offer an elegant graphical perspective on various concepts in the theory of computable functions. In this note, I explore this perspective. (I assume readers are already familiar with Turing machines.)</summary>
</entry>
<entry>
<title>Expectiles are simple to compute</title>
<link href="https://far.in.net/computing-expectiles"/>
<id>https://far.in.net/computing-expectiles</id>
<updated>2025-06-11T00:00:00Z</updated>
<summary>Expectiles are a class of summary statistics generalising the well-known expected value. They have been relatively neglected since their introduction. Perhaps this is because the expectiles of a sample lack a well-known, simple and efficient calculation procedure?</summary>
</entry>
<entry>
<title>Expectiles are configurably-optimistic expectations</title>
<link href="https://far.in.net/expectiles"/>
<id>https://far.in.net/expectiles</id>
<updated>2025-06-11T00:00:00Z</updated>
<summary>Expectiles are a class of summary statistics generalising the well-known expected value [1, 2]. They have been relatively neglected since their introduction [3]. Perhaps this is because they lack a well-known, immediate interpretation like that of the expected value, or that of quantiles?</summary>
</entry>
<entry>
<title>Balanced academic orbit</title>
<link href="https://far.in.net/balanced-academic-orbit"/>
<id>https://far.in.net/balanced-academic-orbit</id>
<updated>2025-05-25T00:00:00Z</updated>
<summary>What is the difference between flight and orbit? Flight is a never-ending fight against gravity. It requires consistent effort to sustain. Given a finite fuel source, landing becomes inevitable. In contrast, orbit uses gravity to its advantage—falling only spurs you forward. Progress becomes essentially inevitable.</summary>
</entry>
<entry>
<title>Attention: All you need to know</title>
<link href="https://far.in.net/attention"/>
<id>https://far.in.net/attention</id>
<updated>2025-03-29T00:00:00Z</updated>
<summary>In this technical note, I explore attention, a key element of the transformer architecture driving recent progress in deep learning applications. I review four standard perspectives on how to understand the attention mechanism. The four perspectives are as follows.</summary>
</entry>
<entry>
<title>Approximate minimax, three ways</title>
<link href="https://far.in.net/approximinimax-three-ways"/>
<id>https://far.in.net/approximinimax-three-ways</id>
<updated>2025-02-25T00:00:00Z</updated>
<summary>This is a short technical note written in the process of working out some definitions for a paper about using an approximate minimax regret training objective to mitigate goal misgeneralisation in advanced deep reinforcement learning systems. I define three different approximate relaxations of the minimax objective, and show that the definitions are related under certain assumptions.</summary>
</entry>
<entry>
<title>Minimax, three ways</title>
<link href="https://far.in.net/minimax-three-ways"/>
<id>https://far.in.net/minimax-three-ways</id>
<updated>2025-02-23T00:00:00Z</updated>
<summary>This is a short technical note written in the process of working out some definitions for a paper about using an approximate minimax regret training objective to mitigate goal misgeneralisation in advanced deep reinforcement learning systems. I define a minimax objective in three different ways, and show that the definitions are equivalent under certain assumptions.</summary>
</entry>
<entry>
<title>Volumetric analysis of the ICLR 2025 review process</title>
<link href="https://far.in.net/iclr2025-volume"/>
<id>https://far.in.net/iclr2025-volume</id>
<updated>2024-12-08T00:00:00Z</updated>
<summary>I was a reviewer for ICLR 2025. I spent a lot of time writing my reviews and engaging with the review process. Much more than (most of) the reviews on my submissions. I was curious about the size of other reviewers’ reviews and also the amount of author–reviewer discussion that goes on during the review period.</summary>
</entry>
<entry>
<title>Transformers are universal in-context function approximators</title>
<link href="https://far.in.net/universal-icl"/>
<id>https://far.in.net/universal-icl</id>
<updated>2024-12-05T00:00:00Z</updated>
<summary>This is a short technical note adapted from a presentation I gave, apropos of nothing, at the metauni SLT seminar on Thursday, December 5th, 2024.</summary>
</entry>
<entry>
<title>Contra Zuckerberg on ‘Open Source AI’</title>
<link href="https://far.in.net/zuckerberg"/>
<id>https://far.in.net/zuckerberg</id>
<updated>2024-07-26T00:00:00Z</updated>
<summary>In internet tradition, I am putting aside my other duties and taking the day to write.</summary>
</entry>
<entry>
<title>Ethics and the Future of Intelligence</title>
<link href="https://far.in.net/future-ethics-lecture"/>
<id>https://far.in.net/future-ethics-lecture</id>
<updated>2024-05-23T00:00:00Z</updated>
<summary>COMP90087 The Ethics of Artificial Intelligence is a Master’s subject at the University of Melbourne designed as an introduction to moral philosophy and applied ethics for students with a technical background. It offers students who will go on to contribute to building AI systems of the future a framework for evaluating, critiquing, and designing AI technology in a socially responsible manner.</summary>
</entry>
<entry>
<title>Safe and responsible AI in Australia</title>
<link href="https://far.in.net/responsibility-letter"/>
<id>https://far.in.net/responsibility-letter</id>
<updated>2023-07-28T00:00:00Z</updated>
<summary>An open letter to the Hon. Ed Husic, MP, Minister for Industry and Science. Submitted through the consultation on supporting responsible AI run by the Department of Industry, Science and Resources from June 2023 to August 2023.</summary>
</entry>
</feed>
