Researchers Stress-Test Alignment Midtraining Across 110-Billion-Parameter AI Models
In a paper published on ArXiv CS.AI, researchers evaluated alignment midtraining up to 110-billion-parameter models and found that its steering effects are fragile against competing finetuning data.
High Signal1 min read
Fragile Steering and Rule Learning
The authors found that alignment midtraining can steer model motivations in simple scenarios, but a tiny fraction of competing finetuning data eliminates those gains. Additionally, models required direct demonstrations in either midtraining or post-training datasets to learn rules robustly. [1]
TechCrunch AI reports that Anthropic is now operating a lab where it conducts biology experiments. The company has positioned AI as a possible tool for curing human disease but also warns of potential risks.
According to a study published on ArXiv CS.AI, padding prompts makes vision-language models more robust to image corruption, whereas fine-grained questions increase model fragility.
Good Start Labs is using strategy games to train artificial intelligence models for real-world tasks, Latent Space reported. In a recent experiment, training a 30B model inside the board game 1830 improved its scores on the Finance-Agent benchmark when configured as a multi-turn terminal agent.