Olivia Moore shares an a16z map identifying search, creative tools and productivity as areas where early consumer AI winners are emerging, while social, marketplaces, travel, finance and other categories remain open.
James Zou announces the Evidence API from Paperclip, which lets users search over 100 million papers to find evidence for and against a scientific claim in about a second. He cites coffee and diabetes risk as an example query.
OpenAI announced it is releasing a range of new mathematical results produced by an internal frontier model, consulting the Institute for Advanced Study's Advisory Group on Mathematics and Artificial Intelligence on how to release them, with materials published on GitHub.
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
Polymarket reports that AI startup BioinvestGPT correctly predicted five of six major drug trial outcomes before results were announced. The post offers no methodology or independent verification.
Economist Noah Smith posts a brief one-line statement that the situation reflects unimaginable levels of corruption. The post provides no specifics about the subject.
Commentator Matthew Berman praises Tesla's Full Self-Driving for his 83-year-old father's independence and says Powershare Home Backup, now available for new Model 3 and Model Y with Powerwall 3, can extend whole-home backup by over two days using a vehicle.
Tesla FSD already turns a car into autonomous transportation. It's one of the most impressive products I've ever used. My 83-year-old father told me he didn't think he'd be able to drive much longer. Then he got Tesla FSD and can maintain independence.
Now your Tesla can also back up your house.
The biggest drawback of battery backup vs. a gas generator has always been limited capacity. This largely solves it.
If your Powerwall runs out after two days, drive to a Supercharger, charge for 30 minutes, come home, and you've got roughly another two days of whole-home backup.
Investor Sam Lessin says he is reading hedge fund quarterly letters for the first time to understand how managers view the current moment. He notes that funds up about 30% in the quarter tend to discuss small businesses.
For the first time i can remember i am reading all the hedge fund quarterly letters because i am truly interested in how they are thinking about the moment... some are quite good. the ones up 30% on the quarter do like taking about small businesses.
Yishan replies to Palmer Luckey's oxytocin remarks with a brief joke titled as reinventing hugs from first principles. The post is lighthearted commentary with no substantive claims.
Palmer Luckey reveals the startup he'd build today: oxytocin doping to end divorce
"I would try to end the tragedy of divorce with oxytocin doping for marriage counseling."
"We understand mammal pair bonding pretty well. There's a lot of people who want to stay together. Most of the people who don't want to stay together should stay together, especially for their kids."
"And we can, if people think SSRIs are fine, which they're not, but if they were, then oxytocin doping to stay together is definitely fine."
"It's the mammal pair bonding chemical. There's things you can do to increase the …
Developer Levelsio reports that many hotels are owned by governments, including Shanghai's municipality via Jin Jiang International's 539 hotels, and says he added a government-owned filter to his hotel-ownership site.
TIL a crazy number of hotels is owned not by public stock holders, not by private owners, not even by private equity funds, but by governments!
You don't see this easily, because you have to dig quite deep (or well dig upward actually) and they cover it up in lots of constructions
Like this hotel in Paris "Le Ballu", sounds French? Yes and it's owned by Hotels & Preference, which on paper is a French international hotel chain
Hotels & Preference itself is owned by the Louvre Hotels Group, sounds French too? Yes
But Louvre Hotels Group is owned by Jin Jiang International, that doesn't sound so French anymore? So it must be a Chinese company? Right?
No! It's the municipality of Shanghai!?
Yes the the local council of Shanghai somehow owns 539 hotels around the world including Raddison, Park Plaza, Vienna House, Golden Tulip, Campanile, Kyriad etc.
Other big government owners are the government of Dubai, Brunei and Qatar
So I added a filter [ Government-owned hotels ] and every hotel page shows the TRUE owner of a hotel, not just the chain!
🔲 Public company ☑️ Privately owned 🔲 Government owned
I love ETFs as much as the next guy but they do put a pressure on stocks to grow by 10%/year or more
Hotels have a hard time doing that, so how do you grow? You cut costs and big hotel chains have been doing that in the name of DEI and ESG and wokeness a lot. Locking down ACs, no daily cleaning, cutting corners, etc. it all saves money
Those cost savings mean higer profit margins mean growth!
Great as an ETF or stock holder, terrible as a hotel guest
The official Grok Bot account announces that it can now search, read and monitor posts on X. The post is a short video announcement with no further details.
Jurrien Timmer writes that strong earnings are offset by a tightening Fed and rising bond yields, with price meandering sideways for four months while market internals weaken. He shares analysis in a LinkedIn post linked from the tweet.
The markets continue to be pulled apart by opposing forces: a massive earnings boom on the one side and a rising cost of capital on the other. In the middle is the residual of the E and the P/E, price, which continues to meander for 4 months now while the market’s internals are bleeding, courtesy of a tightening Fed and rising bond yields. Let’s explore.
Ilya Sutskever posted that people who value intelligence above all other human qualities are going to have a bad time, a short reflective remark on AI and human values.
David Heinemeier Hansson describes a workflow for adversarial code review in which he asks Codex to review work from Claude and vice versa, using each model's command-line interface. He says the models can take turns and settle disagreements without special tooling.
Vijay Iyengar, who helped create Sierra's interview process, responds to points in a post by @jdpruettt on hiring experiments when candidates have capable agents. He says Sierra grades product judgment and system understanding via a two-hour prototype, scope decisions, production thinking, short-answer questions and a debugging interview with a semantic code search skill.
I helped create Sierra's interview process (sierra.ai/blog/the-ai-native-interview), so wanted to reply to some of the points here, which I largely agree with!
Models ship end-to-end demos trivially. What’s the point of testing for this?
We want to evaluate product judgment and system understanding. It’s hard to do this in English, better when you have a real product to look at. The prototype a candidate builds over 2 hours is more a rich visual aid for discussion than the artifact we’re grading. We also find a lot of signal in what scope candidates decide to keep or cut within the time box. The best candidates focus on what makes the product great instead of boilerplate that is trivial to add later. Finally, we spend a good portion of time on how they would take this system to production. We look for whether they really understand what makes the problem nuanced in a real-world setting vs the traditional FAANG style system design where you name-drop consistent hashing and Redis pub/sub to pass.
Short-answer questions
This is something we don’t do in general, outside of certain specialized roles, but I agree it’s valuable. You get a lot of signal about a candidate’s depth in a particular area by whether they can grok and answer questions quickly and concisely. Curious how well this works for generalists vs specialists.
Unassisted code review and tweaking
I really like this idea. We introduced a debugging interview where candidates are given an unfamiliar codebase and have to find/fix a bug in it. We do allow for AI assistance through a skill that doesn’t reveal the problems directly, but acts like a “semantic code search”. We find this to be a happy medium and reasonably representative of the real-world, where you’re just not going to interact with code without an AI anymore.
Doing things that don’t standardize
I very much agree with JD here. The main benefit of Leetcode interviews is that they’re easy to standardize. But imo, you shouldn’t be hiring engineers as quickly anymore, which means it’s worth sacrificing standardization for higher-signal. Of course you want to avoid bias, but I think it’s worth pushing on “what would we do if we didn’t have to standardize?” and questioning whether you really need to double or triple headcount. The three problems with graduated time constraints is a great approach.
Mitchell Hashimoto argues that many highly successful people across industries are very smart, even when they present themselves as unserious, and reacts to a clip of Ben Affleck discussing his machine learning knowledge.
An uncomfortable truth for a lot of people is that a lot of the most successful people in any industry are wicked smart. We apply stereotypes based on the loudest people, but its not representative of norms. Most are crazy smart and got to where they are because of it.
I grew up in LA, one of my parents was a TV producer, my wife was an actor, many friends "in the industry." I've been constantly surrounded by it, and some of the stupidest seeming people (cause its their bit) are actually insanely smart and strategic in private. It sometimes is just beneficial for various reasons to... not show that side of you.
My favorite anecdote was when a friend (a model) was dating a phD in physics and a very rude person made some off the cuff degrading comment to model friend about it without realizing model friend has their own phD in math lol. You wouldn't know from their instagram though!
For the quoted video, I don't know Ben Affleck. I don't know how deep his knowledge goes. He's almost certainly not going to be more knowledgable than someone who professionally does this area of work full time, but he also to me on the surface sounds like someone who knows a LOT more than the average person and possibly even the average software engineer (on this topic).
Anyways, this applies to tech too. I see it constantly.
Ben Affleck reveals he writes Python, understands convolutional neural networks, worked extensively with GPUs, and used his celebrity status to get private looks at Google and OpenAI’s video models
“I’ve always been kind of into computers since I was young. Then, when film started to move from analog film to digital, I became more interested in that aspect of it. The visual-effects workflow for many years has included machine learning, so I can write pretty shitty Python scripts and stuff like that.
“With convolutional neural networks, which were the precursors to what the transformer can do…
Deedy highlights a claimed OpenAI mathematics release, conditional on verification, and argues that LLMs have made major progress on several Millennium Prize problems, concluding that AI has largely reached human-level cognitive ability.
OpenAI’s math release is, in the words of Opus, “the single most consequential mathematical release ever” *
LLMs have how made substantial progress on 4 of 7 Millenium Prize problems: Navier-Stokes (claimed), Riemann, Hodge and Birch-Swinnerton-Dyer. Poincaré was solved in 2003. The two left are P v NP and Yang Mills. *conditional on verification
On average, each OpenAI result used only 3hrs of thinking compute on their new unreleased models.
Two years ago, models said 9.11 > 9.9. Today, we have results that the smartest human minds have not been able to achieve in their entire lives. It is clear that data, compute and algorithms scale. Every model generation (3mos) has made substantial intelligence progress.
It also becomes incredibly hard to not believe that all knowledge work will change monumentally over time. The things AI can not do in the foreseeable future are very likely context-bound (don’t have access to the right information) than intelligence bound. Some might argue they are also creativity-bound or judgement-bound (what should I work on), although it can be argued that future generations could solve for this (given, say, the advancement in research taste for models over time).
AGI is defined as surpassing human capabilities on virtually all cognitive tasks. By most interpretations of that definition, we are there. The domains humans are still better than AI, such as, robotics / physical world control (data-bound?), natural science research (data-bound), some creative domains like writing, movies, music (creativity-bound), super long tasks (context-bound), choosing problems to solve (creativity-bound) and maybe human relationships management (meat-proxy bound?). In many ways, we have achieved AGI.
Daniel Di Martino says a new rule charging $70,000 for OPT and $30,000 for STEP OPT rests on an unrealistic 44% demand drop assumption, and warns it will harm innovation and cost taxpayers.
It assumes demand drops by 44%, which is an arbitrary and incredibly optimistic scenario. There is no indication that universities will pay this, demand will drop 99%.
This is bad for innovation, and will cost taxpayers.
Prajwal Tomar promotes a gallery of motion graphics created with Claude Opus 5.5, each shown alongside the exact prompt that produced it, collected from posts on X and credited to their creators.
Adrià Martinez describes the Claude Startups program, which offers a year of Claude Team, API credits, priority rate limits and tool discounts, and gives step-by-step application tips such as using a company email.
It took me 5 minutes to get into the Claude Startups program
And you don't need VC funding anymore
What you get:
- A year of Claude Team - API credits - Priority rate limits - Deals on the tools a startup runs on
How to apply:
- Create a Claude Console account (the API one, not the app) - Use a company email, not Gmail - Go to claude.com/programs/startups - Fill in the form with what you're building on Claude
Aakash Gupta reports that the FBI arrested Edward Frith after Epic Games reviewed a reported voice clip of his threat and sent it to authorities, and explains how Fortnite's rolling five-minute voice buffer and reporting system work.
The FBI just arrested a Fortnite player using a recording the game made of his own voice. He had no idea his headset was taping him. Almost nobody playing does.
Edward Frith, 29, had logged into his account over 1,700 times. In September he told another player that if the FBI showed up at his door he'd shoot them too. Someone in the lobby pressed the report button. Epic reviewed the clip, sent it to the FBI on September 20, and agents arrested him within days.
The design of the system is the clever part.
Recording every player would be a privacy disaster and a storage bill nobody wants. So Epic built voice reporting in 2023 to work like a flight recorder. The audio buffer lives on your own device, overwrites itself every five minutes, and never leaves your machine unless another player in the match reports it. The moment someone does, the clip gets packaged and sent to Epic's safety team with the speakers tagged.
He said it to one stranger in a lobby. That stranger had a button that turns the last five minutes into a federal exhibit.
I was a producer on Fortnite, and this is the part people outside the building never see. Threats of real-world violence got treated with the same urgency as a revenue outage. The game is full of kids, and Epic acts like it.
Fortnite is actually the conservative version of this. It still requires a human to press the button. Call of Duty has run AI moderation directly on live voice chat since 2023, no report needed. Every major platform with a microphone is converging on the same architecture.
The era where anything said into a gaming headset stayed in the lobby is ending one match at a time.
Amjad Masad says AI-powered reverse engineering and decompilation are advancing rapidly and expects nearly all software to become de facto open source.
Steve Yegge calls Beads the most mature open-source memory system for agents and says it reached 1.5 million downloads, pointing to planned versioned memory features and a proposed wire protocol.
Beads is the oldest and most mature OSS memory system for agents out there. It's the foundation for all my work for the past year, and has matured tremendously with the work of the Gas City folks.
Now that everyone is finally figuring out that they need to "grow" their company brains, I see all these products coming out. You don't need products, you just need Beads. Show it to your agent today.
Mark Suster congratulates Jiake Liu on the launch of Outer Space, which raised $8 million in pre-seed funding led by Upfront Ventures and builds outdoor living spaces that generate and store solar energy.
Developer kyzo says the new product One, a unified inbox and voice interface for talking with AI agents, reached $2 million in annual recurring revenue within five hours. The post links to a promotional video.
a16z reports that Valon raised a $150 million Series D at a $2.3 billion valuation after signing over $200 million in software deals, and describes how it became a regulated servicer before selling its platform to the industry.
.@Valon just raised a $150M Series D at a $2.3B valuation. Within six months of selling software, they signed over $200M in deals.
(And they're hiring!)
How the seven-year-old company got there:
- Mortgage is $13 trillion of consumer debt running on a system built before the internet, and no servicer will trust a new platform. So Valon became one.
- Co-founder Andrew Wang read every federal and state regulation, 18 hours a day for six months, and turned it into code.
- They ran their own servicer on the software until it hit 3x the industry's efficiency, sold that servicer to a larger mortgage company, and now sell the software to everyone else. One of the biggest servicers in the US is moving 4 million loans onto it, nearly 10% of the market.
a16z's Angela Strange sits down with @Valon's Andrew Wang and Linda Du to unpack what it takes to rebuild the infrastructure underneath a $13 trillion mortgage market that still relies heavily on systems designed before the internet.
Linda and Andrew explain why Valon chose the hardest path: becoming a regulated mortgage servicer, translating decades of federal and state regulation into software, and proving the platform on its own loans before selling it to the industry. That foundation made Valon roughly 3x as efficient as traditional servicing and created the system of record it is now usi…
Claire Vo describes using OpenAI's Decisions API with vision to choose the best thumbnail face from 40 minutes of footage for about $0.13. The API, now in public beta, selects models, tools or actions in near real time.
Palmer Luckey says Anduril is investing $3.7 billion in Arsenal-2, which he describes as the largest American shipyard built since World War II, and calls for bold investment in shipbuilding.
Anduril is building America's largest shipyard since WWII. This space needs bold investment and decisive action.
In the 1790s, Baltimoore clippers were the fastest and most advanced sailing ships of their time. King George III didn't stand a chance. Time to run the same play.
Boris Cherny explains his approach to prompting Claude, advising users to give clear goals, specify effort level and verification steps rather than relying on heavy scaffolding.
I am surprised that people are surprised this is how I prompt Claude.
Talk to Claude the way you would a coworker. There's no secret to prompting. There's no need to be overly scaffolded or prescriptive for most tasks -- give Claude a goal, and it will figure it out.
Back in the Sonnet 3.5 days, your prompt mattered a lot. Nowadays, it's much more important to communicate to the model:
1. What you want it to do 2. How much effort you want it to spend 3. How it should verify that it did the right thing
In a quoted interview, Ben Affleck says he writes Python, understands convolutional neural networks and has worked with GPUs, and that he has visited Google and OpenAI to see their video models. The post also quotes Andy Bechtolsheim saying AI has raised optics demand roughly tenfold, with much more growth ahead.
Ben Affleck reveals he writes Python, understands convolutional neural networks, worked extensively with GPUs, and used his celebrity status to get private looks at Google and OpenAI’s video models
“I’ve always been kind of into computers since I was young. Then, when film started to move from analog film to digital, I became more interested in that aspect of it. The visual-effects workflow for many years has included machine learning, so I can write pretty shitty Python scripts and stuff like that.
“With convolutional neural networks, which were the precursors to what the transformer can do, which is much more computation simultaneously, you would do things like look at what’s called a tensor. That’s the numerical translation of a visual image in numbers, like the batch number, the frame number and the red, green and blue values of each pixel in each frame. It’s just that simple. That numeric is called a tensor.
“You’d use a convolutional neural network to identify patterns that reveal what’s called edge detection or feature extraction, which is identifying patterns well enough to know, this is where the window ledge is, so we can more easily take the green-screen image out and replace it with something.
“That was familiar to me early on because, prior to Artists Equity, I had a small visual-effects company. I’ve worked with GPUs a lot too. The visual-effects guys said, ‘Hey, you should see. There are a couple: Google and this other company, OpenAI, are doing really interesting stuff with transformers in video.’
“I’ve learned that I can actually just call up and go, ‘Hey, it’s Ben Affleck. Can I come see what you’re doing?’ Sometimes people say yes, to my astonishment.”
$ANET Andy Bechtolsheim says AI has made optics demand about 10 times bigger in five years and the industry is only at the "very beginning"
"So I guess I don't need to tell you that AI has been driving this incredible increase in demand for high-speed optics, which is probably now 10 times bigger than it used to be five years ago..."
"And what I want to talk to you today is that we're not at the end of this journey, but rather the very beginning. There's easily another order of magnitude increase in bits needed for the next generation kind of data centers."
Eric S. Raymond shares the photocraft GitHub project, a clean-room open-source reimplementation of Photoshop that he says was likely generated by decompiling the app, converting it to a spec and prompting an LLM for Rust. He argues this threatens closed-source software.
This is the doom I predicted a few days ago, coming for Photoshop. A clean-room open-source reimplementation.
No prizes for guessing that they decompiled Photoshop to source code, processed that to some kind of non-code specification language, then fed the spec to an LLM with an instruction to generate Rust.
Adobe just got nuked. And closed source is dead, dead, dead.