Today’s best material is about who gets to judge, and who gets your data. GitHub wrote the test for AI code reviewers, LibreOffice is selling its absence of AI, and an agent startup says trust is its moat while it can send your task to a stranger.
1. Denmark’s ID register reportedly leaks more people than Denmark has
The Register’s headline carries the whole oddity: Denmark’s ID register spilled more people’s details than the country has residents. A leak larger than the population is a data-retention story as much as a security one.
The supplied text for this item is only the headline, so I can’t tell you how the data got out, who held it, or the exact figures. The headline does suggest the register held records beyond the current population, such as people who have died or left. That part is my inference.
If it holds, the lesson for anyone building identity systems is plain. Every record you keep past its usefulness is exposure you carry for nothing.
2. Atlassian warns of a critical file access flaw in Data Center
The Register reports that Atlassian has issued a warning about a critical file access flaw in its datacenter products. The supplied text is only the headline, so I can’t say which products or versions are affected, whether a patch exists, or whether anyone is exploiting it.
The practical consequence is still clear. Self-hosted Jira and Confluence often hold the most sensitive material in a company, including internal docs, tickets and credentials pasted where they shouldn’t be. A critical file access bug there is a drop-everything item for whoever runs those servers.
Read Atlassian’s advisory directly before deciding how urgent it is. This is also a reminder that running it yourself means you own the patch schedule.
3. LibreOffice turns ‘no AI’ into a product feature
LibreOffice’s maker, The Document Foundation, says it “will not add” AI to its software for the foreseeable future. The August release said the software contains no generative AI features, calling that a “deliberate design position” so documents aren’t uploaded to servers and nothing requires a network connection.
The pitch is aimed at compliance. The Foundation argues that for confidential or legally privileged data, the only assurance that survives an audit is that the data never leaves the machine. That is a sturdier argument than “AI is annoying.“
It isn’t a ban. Extensions can connect LibreOffice to local models, and the Foundation says no integration yet meets its requirements, including on-device data and no dependence on one AI provider. So the position is a high bar rather than a rejection. For founders, it’s a reminder that restraint can be a product when the buyer’s problem is liability.
4. Khosla says Wajo wins on trust, but its agent can hire humans
Vinod Khosla told TechCrunch that Wajo’s Fo agent is the most trustworthy personal agent because it was built for trust and safety first. He contrasts it with Meta’s Muse, saying he wouldn’t hand his data to a company in the ad business. That is an investor’s pitch, not a measured result.
Fo does what its rivals do: messaging, errands, phone calls. It also has a virtual credit card that hides your details from sites, and it can hire a human to complete a task. Wajo says it has grown 10x since late September and is used in 107 countries, but gave no user numbers or funding amount.
The reporter’s tests show where trust gets tested. A call to an airline worked, but a call to remind a partner to feed the cats went badly and the disclosures were less clear. The source doesn’t say what a hired human can see, which is the obvious question for a trust-first product.
5. GitHub’s ReviewBench says its offline scores predicted production
GitHub’s strongest claim for ReviewBench isn’t the dataset. It’s that offline results pointed the same way as production. In one Copilot code review test, a multi-model ensemble was predicted to raise precision, recall and comment volume while cutting cost. The A/B test showed addressed rate up 8.0%, recall up 13.6%, comments up 61% and cost per review down 8.0%. ReviewBench predicted a 227% rise in critical comments, against 262% online.
The benchmark uses 219 pull requests from 187 open source repos across 19 languages. Its golden set comes from human reviewers, author follow-up commits, analysis tools and several frontier LLMs. Claude Sonnet 5 acts as the judge. GitHub says senior engineers re-labeled every ground-truth finding and agreed 96.6% of the time.
The caveats are worth keeping. The validation is GitHub’s own, and it also builds a code review product. A benchmark graded and partly seeded by LLMs may favor LLM-style findings. A new entry reaches the leaderboard only if it beats that agent’s current score or is its first. Publishing the dataset, rubric and judge lets others check all of this, and that openness is the point.