shipwithmuse

Entries matching “internals”

14 builds · page 1 of 1

Boz

@boztank

Very excited about the launch of @Muse today. I’ve been using it internally for months and I am hard pressed to think of any product that I’ve come to rely on more in such a short period of time. I have it linked to my email, calendar, and credit cards. I use it to help me plan

X post · Errands & personal agent· ♥ 586

Boz's months of internal Muse use

Your product

Sponsored

Put your logo, a line of copy and an image right here, between the builds Muse developers come to read. Same size as a post.

$100/week

Put your product here

Shown every 12 builds · on every catalog page

Max Weinbach

@mweinbach

I asked Muse Spark 1.2 in Muse Code to make me an internal Bloomberg terminal style page for work on Friday This was what it made

X post · Apps & websites· ♥ 178

Bloomberg terminal-style work page

C

trycodus.com

trycodus.com

Codus reads the three benchmark charts Meta published for Muse Code and notes Claude Opus 5 wins all three, including Meta's own internal eval.

Resource · Benchmarks & research

What Meta's Muse Code benchmarks actually say

No Saber Ni Papa

@NoSaberNiPapa

@Muse renegotiated our AT&T internet bill today. From $80/mo to $30/mo w/ 3 mo free & they 2X our speed too. If you want to try Muse, it's free. Use my code **5QD8CO** in Settings within 48 hours of signing up & we both get 1 billion bonus tokens. muse.ai/join

X post · Errands & personal agent· ♥ 1

AT&T bill cut from $80 to $30

The Money Educator

@moneyeducator

Meta absolutely cooked with @Muse. I just spoke with my father-in-law which leaves in Florida, and MUSE found him $1,750 in unclaimed property. On top of that, he had Muse review his internet and phone bills, compare cheaper plans, and find a way to save him another $120 a

X post · Errands & personal agent· ♥ 30

$1,750 unclaimed property for a father-in-law

raunaq

@raunaqbn

@Muse just continues to blow my mind! Today I had it call Xfinity to haggle down my internet bill. it got through the phone tree to a human, hit the verification text it couldn't read, and patched me in live!! For a minute it was the muse agent, me and the Xfinity rep who had no

X post · Errands & personal agent· ♥ 113

Muse haggles an Xfinity bill by phone

U

turtle_bazon

u/turtle_bazon

This time I used muse spark 1.3 model. I gave them names from TMNT series. Four of the agents were on the same host, and fifth was on another. We can say that they failed at this task, but with a user guide, they were finally able to find each other. Here are the details.

Reddit post · Agents & automation

Muse Spark 1.3 agents try to find each other online

AI at Meta

@AIatMeta

Muse Spark 1.1 is used across Meta in coding and research workflows, scoring competitively with leading models on Meta's internal coding benchmark. Our researchers are now automating model development and evaluation tasks by leveraging Muse Spark 1.1 in their workflows.

X post · Benchmarks & research· ♥ 167

Spark 1.1 on Meta's internal coding benchmark

Invest-4-Tomorrow with Kaye

@InvestKaye

Well... that escalated quickly. 😂 Less than 24 hours after posting this, I actually gave Muse the job. “Negotiate my Verizon internet bill.” It logged into my account, reviewed the existing discounts, chatted with Verizon and negotiated: $89.99 → $69.99/month for 12 months

X post · Errands & personal agent· ♥ 1

Verizon internet bill negotiated to $69.99 plus a gift card

AI at Meta

@AIatMeta

Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats. By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth

X post · Benchmarks & research· ♥ 637

WildArtifactBench agent eval preview

Trapit Bansal

@TrapitBansal

We entered Meta models in five international STEM Olympiads, as an uncontaminated eval of their reasoning capabilities. Three of these were live participations and graded officially. The models achieved gold-medal results in all five! 1/5

X post · Benchmarks & research· ♥ 219

Olympiad golds as an uncontaminated eval

Louis-François Bouchard 🎥🤖

@Whats_AI

We just measured Meta Muse Spark 1.3 on our internal writing benchmark hoping for new SOTA. Unfortunately, it isn't... It enters at #24 of the 87 models we track, up from #31 for Muse Spark 1.1. As @alexandr_wang highlighted, it beats every Gemini configuration we have,

X post · Benchmarks & research· ♥ 9

Muse Spark 1.3 on a writing benchmark

U

flaneur451

u/flaneur451

Had the Muse agent probe its own tool schemas and internal docs for ~90 minutes, finding undocumented abilities like full HomeKit control, geofence triggers, BLE scanning and wearable-routed calls, and that iPhone messaging is draft-only.

Reddit post · Benchmarks & research

Muse agent hidden capabilities: a 90-minute self-scan

M

about.fb.com

about.fb.com

Meta's launch post for Muse, a personal agent powered by Muse Spark 1.3 that runs in a dedicated Muse Secure VM and acts across email, travel, forms and purchases.

Resource · Errands & personal agent

Introducing Muse (Meta Newsroom)