Cruxes for alignment

AI alignment is far from a solved problem. We need only look at recent events to see this. Before we can attempt to propose a solution, we must first get the basics straight: what does AI alignment even mean?

Here’s an exercise: what do you think AI alignment is?

I probed my own understanding and realized that it’s really easy to make assumptions or statements without having a clearly articulated model. So, I set out to consider what AI alignment might mean, given different cruxes.

As a starting point, I found it helpful to diagram these scenarios. There are many possible dimensions or frames for considering this problem, and I chose some salient axes to project onto.

There is much more to be done. I will be engaging with prior art and thought experiments to develop and refine the model presented here. It is by no means complete at this time.

Pasted image 20260808231036.png

A note on the schema: colors reflect my confidence in a crux. Green = decently confident, yellow = unsure.