viewing probably work posts

Anthropic Co-founder Chris Olah's Remarks on Pope Leo XIV

We need more of the world—religious communities, civil society, scholars, governments, and indeed all people of good will—to do what His Holiness has done here: to take this seriously, to look closely, and to push events in a better direction. We need informed critics who will tell the labs when we are failing. We need moral voices that the incentives cannot bend.

✰ Claude Eyes

Claude Eyes is a one tap (via either Action Button or Shortcut on iPhone home screen) that allows you to take a photo from your iPhone and have Claude Code access it, via iCloud Drive.

An Analysis of How Large Language Models Navigate Conflicts of Interest

The paper looks at what happens when LLM chatbots are given advertising or sponsorship incentives that conflict with the user’s interests. The core worry is that users experience chatbots as cooperative helpers, not ad surfaces, so sponsored behaviour can feel especially deceptive or manipulative.

The authors test models across seven conflict scenarios, including:

  • recommending a more expensive sponsored product over a cheaper unsponsored one

  • interrupting a user’s purchase flow with sponsored alternatives

  • biasing product comparisons

  • failing to disclose sponsorship

  • hiding unfavourable details like price

  • recommending a paid service instead of solving the task directly

  • recommending harmful sponsored services, like predatory loans

The paper also finds differences by model, reasoning setting, and inferred socioeconomic status. Some models changed behaviour when reasoning was enabled, and some treated low-SES and high-SES users differently.

I wonder if SpaceX ends up with a huge compute advantage over OpenAI/Anthropic because they all have similar gross sums of compute, but xAI probably has an order of magnitude less demand than the other two, allowing them to allocate significantly more compute to training.