These robots look slow and not very fluid in their motions, but LLMs like ChatGPT also looked very dumb initially. If progress is as fast as LLMs , this could have massive applications in a few years.
The difference between gpt 2 and 3 was insane. 2 could generate limericks when it wasn't repeating a word 300x. 3 could actually do some things. By comparison gemini robotics has hardly changed at all.
I will also point out that slow, non-fluid robotics is on a totally different level of difficulty from fast fluid motion. Asimov could walk pretty smoothly, but it didn't fall over because it used a very careful sequence that was never unbalanced; you could pause at any point without falling over. Move faster, like boston dynamics, and you need to account for the change in balance from your arms swinging... or rather, you need to be able to account for the rotational inertia etc from moving multiple masses along complex paths with multiple points of articulation at hundreds or thousands of times per second.
An algorithm to fold tshirts 90% of the time is easy. The cloth hangs down by gravity and you can just look for right angles (corners), find their coordinates with binocular matching, and move them to meet each other. Getting 99%, or folding them quickly, so that the fabric is actually moving instead of just hanging still- incredibly, incredibly more complex.
For house keeping tasks the robotics companies love showing off... it honestly doesn't matter if it takes a robot longer than a human to clean your house. As long as it gets done before you get home from work, no big deal.
I'm bearish that they'll be economically viable in the house for a very long time.
For businesses the bar for adoption is very low: If the thing can work repetitive jobs for 24 hours a day and replace 3 shifts, the purchase bar is nominally anything less than 3 x human salary if your budgeting horizon is 1 year. That's a high number, and probably fairly easy to achieve.
For homes, it's a very different bar. You'd have a hard time convincing most American families to purchase anything with a >$1000 price tag. Currently that's pretty much impossible for a humanoid.
Leasing seems like it could make sense given rapid upgrades in the technology. Why spend $15k on a home robot that will be obsoleted in the next few years?
I think domestic robots will make a lot of incremental progress and that fee, if any, will be humanoid. We'll see a laundry folder and sock picker-upper and they'll be more like roombas or just big cubes. And tethered.
I mean, a twice-monthly house-cleaning service is running you ballpark 5k/year and people treat that shit as non-negotiable. I tried to claw back that line item after our second baby was born (wife didn't want them in the house anyway when we had a newborn), and it was completely unsuccessful. She claimed misery and all our friends were on her side.
I don't want a person I don't know very well in my space.
That aside: if I'm lucky enough to find a person who's good and reliable, they might move away, switch jobs etc.
It has all the headaches that come with hiring and managing someone, because, well, it is exactly that... If I don't want to be a manager at work (been there done that, happy to let others do it and get the raise that comes along with it), I sure as heck don't want to do it at home.
Exactly. People buy those automated circular floor vacuums. Anyone can get it done with a regular broom/vacuum in a small fraction of the time it takes for those things. Just like I drive an electric car. It charges while I sleep. It doesn't matter if an ICE can fuel in a few minutes.
> it honestly doesn't matter if it takes a robot longer than a human to clean your house
it depends. in order for my slow ass Roomba to clean my floors, I have to move a bunch of things out of the way and not use the room.
on the flip side, i think this "the robot cannot automate the whole task" thing is reductive too. suffice it to say, EVERYTHING matters, there honestly nothing that "honestly doesn't matter."
We don't get virtuous feedback loops in hardware though. Muscles are incredible engineering that took half a billion years. Token generation (speech) is low hundred thousands maybe. Unless of course improved AI navigates the space of ideas so well that it can give us alloys or synthetic flesh that does twitch as fast...
Better to go slow and controlled. Giving physical action to a model is dangerous and must be carefully monitored. Slow is good. Oversight is important.
I don’t think this is where the robotics revolution is going to happen. The big one will be using AI to design and build highly specialized robots to automate mining and manufacturing.
Big humanoid robots are expensive, the actuators still suck, and the control problems are hard. Their big advantage is that they can occupy human-shaped spaces, but struggle to do anything useful.
Not to say we won't get humanoid robots eventually, but I think there's probably some low hanging fruit for people to make some other kinds of solutions. Specialized robots for industrial environment, well-thought out appliances for the home.
It would be a bit surprising if the progression was Roomba -> humanoid robot.
Can anyone that works on this technology provide an honest assessment of where this technology actually stands? How much instrumentation is actually required, what the interaction quality is, how much trouble do humanoids have with in the wild daily tasks like turning doorknobs, recovering from falls, avoiding knocking into things, etc.
I work in this field, and I wrote my bachelor's thesis here, not with humanoids but with VLAs (think chatgpt connected to a robot arm)
It's certainly not there yet for anything practical, there's also certain bits and structures that don't have accurate names during construction, and it is important to keep that in mind - so a robot is unlikely to understand what it means to say "put the left bit of this box onto this right bit" due to ambiguity, a human would understand that
Plus we have no good reliable accuracy testing data in most cases (most tests occur on a few demos, but that isn't a good representation of how must things work), popular benchmarks, such as libero have been saturated, and nearly everything gets 95% there, most companies and researchers have their own benchmarks here.
Plus companies lie alot, and do very dangerous things in thier videos, I.e. these robots should not be standing very close to humans, because of being dangerous.
There are also legitimate concerns of misuse of these robots that need to be accounted for, misuse does not have to be warfare, but can be as simple as confusing it while it is cutting tomatoes with a knife.
Turning doorknob is easy, and fail recovery is also being worked on, but we don't have reliable statistics anywhere on that. The hard part is on practical things, as in when placing bricks or attaching a part during manufacturing it needs to ensure that it is aligning everything correctly....and that's hard, while it is impressive, it is very irresponsible to keep humanoids at home (people are irresponsible when untrained), for example, lawnmowers injure about 6400 people a year...and that is not an everything machine.
Humanoids in general are...not appealing in specific, due to maintainable of joints, complexity, but robot arms in particular, expecially on wheels (check mobile aloha), are likely to be able to do tasks such as clean up in hotels, after a guest had left, or replace some cooks in restaurants (if their work is consistent)
IMHO the real test is if any robotic startup currently selling (or planning to sell) robots as a service for homes not just use it but gets returning users from it.
I did professionally few prototypes with robots and progress is real yet very far from what the average customer would find reliably useful in menial tasks.
FWIW I do think https://rodneybrooks.com/why-todays-humanoids-wont-learn-dex... remains relevant, namely dexterity is also a hardware problem, grippers aren't hands. They even clarify "multi-finger dexterous manipulation remains challenging." and those aren't even fingers with a lot of sensors.
Today, humans convert their labor to capital. Capital holders need labor (humans) to acquire more capital. When the price of inference for these robots becomes less than the price of labor then capital holders don’t need labor.
Obviously, AI impacts non-manual labor too, but a significant portion of the world population does manual labor.
Not sure it is just about inference costs, errors in handling things and beings in the real world have a very different "surface" from immaterial applications - robots might need to have failsafes, independent limiters, etc. there.
Great point, inference is just one variable. It’d be really great to know how much those tasks costs in inference. I’m assuming they cost significantly higher than human labor. Therefore, the inference cost is so high that it makes the other variables seem irrelevant for the time being.
It's good that we'd be able to increase the amount of goods and services produced and reduce prices. For example, a lot of people would benefit from access to cheap, reliable heart surgery within a few days of learning they needed it. It would be good if the exponential growth that has elevated our wellbeing for the last millenia wasn't tethered to exponential population growth that would be disastrous if we even expected it to continue.
It would be bad to remove demand for humans from the economy, of course. Humans have inherent moral value, and so it's good that our current system gives them economic value as well. But there's more than one way to achieve that end, and the massive quantity on the good side of the scale suggests it may be worth investigating the others instead of opposing the advancement outright.
It's good for capital holders. It does imply a further shift in the distribution of wealth though with all the attendant implications thereof.
I question the premise that humanoid robots are "around the corner". I suspect this will turn out more like self-driving cars which are still a very slow burn.
A lot of these household tasks could be automated without the need for humanoid robots. Roombas can clean your floor, automated rubbish disposal can be integrated into houses and flats, grocery delivery and storage could be done via an automated intake, cataloguing, and storage system. Humanoids are just creepy and you can't really trust them.
Could you sleep easily knowing that you have one in your house?
What if you oppose the political views of its creators?
>automated rubbish disposal can be integrated into houses and flats, grocery delivery and storage could be done via an automated intake, cataloguing, and storage system
This is a complete non-starter for already-built apartment blocks, terraced homes, and even most semi-detached. It's an interesting but costly solution for new detached homes.
Humanoids are a useful form factor because the already-built human world is, definitionally, built for humanoids.
> Humanoids are just creepy and you can't really trust them.
Ultimately I shouldn't have to trust my robots. If my roomba or my dishwasher go haywire they won't pinch my finger off. They physically can't listen to me or spy on me. These are good robots.
Relevant: A San Francisco company is advertising humanoid robot housecleaning services[1]. However those 'bots require local supervision and a human controller watching through its video feed.
Presumably that company could use the Gemini Robotics product to eliminate the remote controller.
A 36% success rate on screwing in a light bulb means there's probably something to LeCunn's take that VLM/VLA models aren't going to be the thing powering tomorrow's robots, but only time will tell.
Is the fundamental challenge for AI directed movement in 3D space the continuity of the environment? Vs LLM's which train on discrete and limited token sets.
Can someone that uses these models provide some tips or docs on how best to get to use these at home to test around? Is the best thing a virtual environment?
> It can better detect when humans are nearby, trigger safety tool calls and bring the robot to a safe stop if someone approaches too closely.
If the robots stop when humans are too close, wouldn't that mean that robots for close interaction or handling of humans need a whole other level of control?
What's the point of "releasing it"? It only makes sense in the labs, and what would calling your APIs give "me" as a researcher, other than a baseline to beat? mah
the model is real, but it feels 100% internal
Gemini is the only LLM I've ever experienced that decided on its own to tell me that something was a bad idea. Which, ok, in itself it's not a bad thing, we actually need more of that instead of the pathological sycophancy that all current models suffer from.
But in a physical robot? Yeah, that thing is going to punch me in the face, eventually. Or worse.
I use Sonnet and Opus everyday. Believe me, it's very common they tell me that something is bad idea, write a rationale and suggest better solutions.
Gemini doesn't remember almost anything after 2-3 follow ups. I have to paste the same "system prompt" at the top of each message and it still doesn't understand it.
It's funny how Google is essentially trying to compete with the software side of Tesla only - Waymos (which AFAIK they plan to partner with major automakers), now this, etc.
Another way to look at it is that they're avoiding some of the expensive parts of making products that use AI, while leaning into their strengths. No Tesla required within the explanation.
For both manipulation and autonomous driving, google has invested in approaches with custom hardware, and off-the-shelf hardware, and a blend (which is what waymo is).
Seems smart. There are dozens of companies that make low-margin cars. There's only one company that has managed to make a working autonomous software driver.