• unwind ai
  • Posts
  • Manus personal agents with phones, wallets, and computers

Manus personal agents with phones, wallets, and computers

+ Anthropic releases Claude Sonnet 5.5

In partnership with

Start here ↓

Personal agents now come with phones, wallets, computers, and human operators

This week's personal-agent launches are choosing different superpowers. Grok Bot is pushing AI teammates. Muse is pushing a proactive agent that knows you. Today is Manus 2.0, giving agents identity, tools, and a persistent place to work.

The new Manus app, Cue, gives each agent its own email, phone number, wallet, and computer. It can take calls, send messages, make approved payments, and keep working after you log off.

Put multiple agents into a group chat, give them one goal, and let them split the work across research, shortlisting, outreach, and deliverables.

Manus 2.0 is live on web, desktop, and mobile. Cue is free in early access with invite code MEETCUE, with the iOS app still waiting on App Store approval.

🚀 Shipped

Agents are now 48% of Cloudflare's CLI traffic, so Cloudflare wrote the CLI for them. Cloudflare released cf, an open beta CLI that exposes 3,000+ Cloudflare API operations instead of Wrangler's roughly 280 command paths. cf is open-source, defaults to JSON, adds command search, and supports TypeScript config. Install the beta with npm i -g cf and migrate an existing Worker with cf migrate.
Cloudflare

Sonnet 5.5 is the cheaper and faster daily driver. Claude Sonnet 5.5 is here as the faster, lower-cost complement to Opus 5.5. Pricing stays at Sonnet 5 levels, but the model generates 30%+ faster and can cost up to 30% less per task because it uses fewer tokens. It scored 70.6% on Terminal-Bench 4.0 versus 10.3% for Sonnet 5.
Anthropic blog 

Give every agent a real VM. boxd is selling persistent KVM Linux machines that boot in under 10ms, fork a live running machine in under 200ms, and hibernate when network activity stops. Give each agent, branch, sandbox, or developer a full machine with root, systemd, Docker, 100GB of disk, and persistent processes.
boxd

This personal agent hires humans when software gets stuck 💀 Ex-Google DeepMind-er launched Fo, agents that come with their own inbox, phone number, voice, and credit card, and the builder that lets anyone create one. Fo can even route work to human experts when tasks need real-world judgment or phone calls. Per their internal evals, Fo completed 71% of tasks with a trust rate of 94%, way ahead of Muse, Instinct, OpenClaw, Hermes, and more. It is invite-only alpha, with free unlimited access for the first 1,000 users.
Wajo announcement

ElevenLabs makes voice agents whisper, laugh, and respond in 150ms. They shipped Eleven v4 and v4 Turbo for expressive text-to-speech in agents, games, ads, dubbing, and long-form audio. The model follows inline direction like laughs, whispers, pauses, accents, rain, or phone buzzes, so you can steer delivery without rebuilding the whole audio stack.
ElevenLabs blog

Claude and Codex can now use a real iPhone. Tapkit is a Mac app that lets an agent use a physical iPhone through screenshots, taps, and typing. You can drive it with Tapkit's built-in agent, Claude, Codex, any MCP client, or the REST API. Nothing gets installed on the phone, and there is no jailbreak, developer mode, or simulator requirement. The setup constraint is physical: the iPhone needs to be plugged into a Mac running Tapkit.
Tapkit

Cloudflare open-sourced the generator behind cf. Forge is Cloudflare's open-source pipeline for generating SDKs, CLIs, docs, examples, and future surfaces like MCP servers from API definitions. It runs in CI and creates preview builds when API teams change schemas, so teams can test generated outputs before shipping.
Cloudflare Forge

Fireworks’ router kept 98.1% of Opus accuracy for 57% less cost. They launched FireRouter with Opus, a cache-aware router for coding workflows that chooses between Claude Opus 5.5, GLM 5.3, and GLM 5.3 Flash per user turn. Their internal A/B test cut coding-session cost from $15.36 to $6.63 while retaining 98.1% of Opus-only accuracy. Works through Fireworks CLI integrations and as a serverless router model.
FireRouter

🧠 Worth Knowing

GPT Researcher dropped embeddings for Jev. It now uses Jev as its default context filter when a TypeSafe key is present. Instead of embedding similarity, it scores whether each chunk is useful for the research question. Their benchmarks show that Jev kept 73% of relevant passages versus 46% for embeddings at the same cost, with keyword fallback when Jev is unavailable.
GPT Researcher docs

Claude Code can build an eval, then hillclimb your app against it. Anthropic's Claude Devs team added /claude-api build-eval and /claude-api hillclimb to the claude-api skill. The first command interviews you, samples production-like cases, proposes a grader, and runs a baseline. The second improves the app one change at a time against a held-out set, with checks for variance, overfitting, impossible tasks, and broken graders.
Claude blog

Glean put Jev through four real enterprise jobs. Tony Gentilcore at Glean ran Jev against four bounded enterprise decision tasks: query classification, model routing, reranking, and citation-support judging. Glean's routing test showed an 8.1x median speedup on positive transfer cases, while reranking still needed production tie-breaks.
Tony Gentilcore's blog

OpenAI DevDay starts at 10 AM PT today. It is today in San Francisco, with Sam Altman opening the keynote at 10 AM Pacific. Expect technical sessions, demos, workshops, and API/tool programming, plus DevDay Exchanges coming to Bengaluru, Tokyo, Seoul, Paris, Berlin, London, São Paulo, and Mexico City.
OpenAI DevDay 2026

Run seven tiny LLMs in your browser, with no server or account. MicroLLM Lab runs Q4 tiny language models directly in the browser through WebGPU. The catalog includes 26M to 360M-class models, mini-benchmarks, speed tests, and more. It is a playground, not for production, and perfect for feeling what tiny on-device models can and cannot do.
MicroLLM Lab

Run a Jev-compatible decision model locally with Unsloth. Unsloth added docs for serving Laya, an open Jev alternative, from the local Unsloth app. You can turn on the Decision API and point existing TypeSafe-style integrations at localhost for choices, yes/no probabilities, and scores.
Unsloth docs

Anthropic's IPO leak shows $42B loss in 2025. Reuters says it reviewed Anthropic's confidential IPO prospectus and reported a near $42B 2025 net loss, including a roughly $34B accounting charge tied to financing value. It’s a confidential draft S-1, not public yet.
Reuters

🔧 Clone and Run

Clone & Run of the Day

Turn Claude into a daily buyer-signal monitor with Treg leads-signal. Describe your buyer once, then have an agent check hiring, funding, tech-stack changes, job moves, LinkedIn engagers, Reddit, X, GitHub, and reviews on a schedule. It’s an open-source Skill you can use with any agent.

Run Jev-style decisions locally with Jeff. It has the weights for tiny Qwen3.5 and Gemma 4 fine-tunes that return calibrated option probabilities in a Jev-compatible format. Use it for local routing, tagging, and classification. It is English, text-only, and weaker on reasoning-heavy tasks.

Cua Driver can now see canvases, games, and remote desktops with Cua Perception. Cua Driver lets Claude, Codex, or another agent control a real macOS, Windows, or Linux machine, normally through its accessibility tree. The optional extension turns a retained screenshot into typed text and icon regions when accessibility trees fail.

Experiment with System-One-controlled agent memory using Jev-Mem. It uses Jev-style Noul and Choice decisions to organize memory, enforce write policies, and control retrieval. It is research code with demos, benchmarks, and tests.

Expose a Jev-compatible decision API from llama.cpp models with jeva.cpp. The fork adds Choice, Score, and Noul endpoints to llama-server while preserving normal generation. It should work across llama.cpp-supported models and backends, but the decision API needs model types that expose next-token vocabulary logits.

Make launch-style business videos directly with Claude with Motion Video Kit skill. The repo packages a Claude Code skill kit with motion rules, critic loops, audio checks. The builder learned from studying 28 professional SaaS launch films and from building two full sample commercials through dozens of rounds of independent critique.

Awesome LLM Apps is a curated collection of 100+ AI Agents, Agent skills, and RAG apps. It covers models from OpenAI, Anthropic, Google, and open-source models like GLM, DeepSeek, and Qwen that you can run locally on your computer. (Now accepting GitHub sponsorships)

That's all for today. Come back tomorrow for the next batch of AI tools, model drops, agent repos, and weird benchmarks worth your time.

If you found one thing to try, share the issue with someone who ships.

Unwind AI - X | LinkedIn | Threads

The ice cream shop that makes money when it's cold

28 Wishes sells ice cream in Los Angeles. Below 70°F, sales fall about 20%. The weather is out of their hands. Rent isn't.

So the owners started putting about $20 a day into Kalshi weather markets, taking the cold side. The days that keep customers away now pay something back.

This is hedging. Big companies have done it for decades, buying protection against bad weather, fuel spikes and rising rates. It used to take a broker, a trading desk, and an order size no corner shop could meet.

Kalshi opens it up. Contracts on weather, fuel prices, inflation, tariffs and regulation, starting at a few dollars. Take a position on the outcome that would hurt you. If it hits, the payout softens it. If it doesn't, the contract expires and the good month was the point.

Reply

or to participate.