Mathematical Scientist for AI Safety Research
About the Role
Frontier AI companies are throwing billions of dollars into scaling existing architectures and methods such as next-token-prediction, direct preference optimization (DPO), reinforcement learning with human feedback (RLHF), and reinforcement learning with verified rewards (RLVR). These methods are very powerful, yet fundamentally flawed, resulting in misalignment, sycophancy, systematic biases, and other forms of harmful behavior that are already having severely negative consequences in our society.
LawZero is a…
Unlock the apply link on every AI job
Browsing is free. Membership is $9 a year or $4.90 every three months and unlocks the apply link and the full description on every job. Cancel any time from your billing page.
- The apply link on every job, straight to the official posting
- Full job descriptions instead of the preview
- Every AI job, every day: thousands of listings from 1,000+ company career pages
- Save jobs and get daily alerts for your categories
Secure checkout by Dodo Payments·Cancel anytime from your billing page