People really need to stop thinking about 5% yields over 10 years or 30 years.
Even the inflation of my $SUBWAY sandwich has gone up like 12% a year.
Today it was like $20 after tax, when it used to be like $5… 12 years ago.
The best way to keep up with inflation of the things that matter, like Sandwiches, is equities like $SPY.
投资者运行 1 万个 AI 工作负载,测试 Meta 免费 Muse Agent 的成本压力
投资者 Lee Roach 透露,在其团队将 Meta 的个人 AI Agent Muse 扩展到 1 万个持续运行、满负载的工作实例之后,他建立了规模较大的 Meta Platforms 空头仓位。
Muse 于 9 月 8 日推出,可执行发送邮件、预订旅行等任务,目前在美国通过 iOS、Android、网页端和 WhatsApp 免费提供。为了吸引用户,Meta 承担 Muse 每次查询所产生的计算成本。
Roach 表示,每部署一个 Muse 实例,他这边只需要承担大约 300 美元的一次性前期成本,但此后持续运行产生的、理论上没有上限的经常性计算费用则由 Meta 承担。
基于这一点,他认为,如果大量用户以类似方式高强度使用 Muse,可能会给 Meta 的利润率带来压力,甚至导致未来盈利低于市场预期。
这一做法也引发了争议:一些人认为,这是对 Meta AI 产品商业模式和单位经济性的一个聪明的“压力测试”;另一些人则认为,这更像是为了从做空获利而有意制造高额资源消耗,尤其是在 Meta 正投入巨资建设 AI 数据中心的背景下。
Railroad construction in the early 1880s peak near 6%
In addition, the railroad investment excludes rolling stock but the AI number includes 68% computing and networking
@bubbleboi I’m not very optimistic about making a good Iran deal in the short term. And I think the market might already priced in a lot of it.
Also I think the current bottleneck is the refinery not crude oil.
@Z04_A1 Dollar bulls are crowded, bond bears are crowded, Russel bears are crowded.
We’re in a favorable environment for bulls and need a catalyst to ignite the rally.
I don’t know what it’d be and when it’d be. So I chose leap calls
And when the tide turns, I can chase quickly.
I’ve further reduced the exposure by converting into leap calls today and entered the standby mode until Oct FOMC.
Also holding few short dated calls to bet OAI dev day next week.
I think bond bears might cover shorts before PCE to give the equity market some relief.
Be Patient!
@btc__gabriel @kayliatyyy I think if they manufacture and sell in US, they have to follow the US rules and privacy standards , and the US company must be a co-owned structure with US shareholders taking the main control of the board. it would be similar to what TikTok is doing and its structure now
Bigger AI Chip Packages Create Bigger Materials Challenges
Economic Daily News reports that larger packages and more complex chip stacks are increasing warpage, heat-management and reliability challenges. That raises demand for underfills, encapsulants, redistribution-layer materials and cleaning chemicals as CoWoS and WMCM production expands.
Taiwanese chemical suppliers now have an opening to localize more of the advanced-packaging supply chain.
Liwei’s take: Taiwan’s materials moat is more than chemistry. Dense supplier clusters, fast on-site troubleshooting and years of customer trust help turn process problems into qualified solutions faster. Subsidies can build a fab; rebuilding that working network takes time.
Sources
https://t.co/SxFrCkruHo
Brilliant Analysis from @harry03994688 : also $AVGO is a natural rival to $NVDA.
Its ASIC business directly competes for the same AI compute dollars—and that conflict may also be costing Broadcom opportunities elsewhere in the NVIDIA ecosystem.
Totally agree. While some creators naturally lean toward a more text-heavy style where the main point can sometimes get lost, those who follow our small circle know we like to keep things simple.
Just one word:
testing, testing, testing.
and just in case anyone missed it, we repeated it three times:
Testing, testing, testing.
How much alpha is packed into that single word? We'll let you figure that out. 🤫
Oh, by the way, looks like @RYANHINGSHING , @ShanghaoJin Jim, and Liwei’s public holding, $AEHR, is already up 40%+ in September alone.
What survives training-data cleanup may shape the value of adding more Engram memory.
It was also designed with serving in mind, using token IDs mean retrieval can be async and overlapped, allowing DRAM offload without performance loss. (5/6) https://t.co/KPb8aVuMiP
Copyright notices, licensing text and sharing prompts also show up.
Engram learns what helps predict text, not what humans think deserves remembering. Even boring web boilerplate can offer useful shortcuts. (4/6) https://t.co/kHGrZsEoPb
Layer 1: “pints of frozen yogurt”
Layer 14: “three times as many”
The first describes the objects. The second expresses a reusable relationship. Different layers can use memory differently. (3/6)
Engram retrieves learned vectors for short token sequences, helping the model reuse familiar patterns instead of reconstructing them.
Our scan found names like “Ian Goodfellow”, code fragments, task instructions and website boilerplate. (2/6)
DeepSeek’s memory lights up for "Wright : Ace Attorney"
We probed V4.1 Flash’s Engram gates to see which text patterns it uses. The results go well beyond names and facts. (1/6)🧵 https://t.co/EplmijW3d9
The SemiAnalysis STEEL teardown lab is hiring! Are you driven to explore the technical depths and nuances of advanced semiconductor manufacturing and design? Love a good floorplan and understand how it all goes together? Join a killer team. Apply today at (3/3) https://t.co/uydSzy75Cz
SemiAnalysis STEEL teardown lab has created something big. We're tearing down advanced datacenter and AI hardware and for fun we're sharing consumer teardowns for free! To learn more about our pipeline or to commission a teardown, contact sales@semianalysis.com (2/3)
Tearing down Apple M6 and TSMC N2. We're sharing it all here, free. TSMC makes three (GAAFET foundries), scaling, optimization, and a few surprises. Stay tuned! (1/3) 🧵 https://t.co/nvTpaICoWR
It is worth noting that our hands-on testing is only a part of the overall ratings. There are many things that we cannot test hands on, and need to use other research methods to understand in detail. Namely, performance at scale, reliability over time, support experience, pricing, GPU availability and delivery timelines.
Read more in ClusterMAX 3.0 (7/7)
https://t.co/hX3qLA6fR8
This is our most critical test. Providers where we hear a bunch of customer complaints about reliability generally do not have health checks, monitoring dashboards, and autoremediation in place. Reliability is the #1 most important criteria to many of the biggest customers in the world, as we discussed in great detail in our article. (6/7)
https://t.co/QsZgMrKRl1
For injecting failures, we start by using the DCGM injection method. If it doesn’t work properly, we follow up by writing a synthetic NVIDIA XID or SXID message into the kernel log, and watch how the provider’s existing health checks and scheduler respond. We leave their health agents and drain automation alone, so it’s up to their tooling to detect the error and act on it. Most health checks are set up to read from the kernel ring buffer, but some are not, so we communicate with the provider to make sure that our trigger will work and that we have access to pull it. As a final, trusty method, we reset the GPU’s upstream PCIe bridge, which produces a genuine XID 79. (5/7)
First, we reboot all the nodes in the cluster. You would hope that this is not a destructive test, but it is. On Kubernetes, we cordon and drain the node, issue the reboot, and require a changed boot ID, then wait for it to return to the cluster. The timer stops when we can run nvidia-smi in a fresh pod on the node. With Slurm, we just get an allocation or SSH and run sudo reboot. Slurm-on-Kubernetes is a little different, so we attempt to do things from the Kubernetes layer to test it correctly. Our scripts time everything throughout. (4/7)
Next, we test the fabric, storage, and orchestration software individually with specific tests designed to measure the sustained performance under load. Finally, we break things on purpose (or simulate them breaking). (3/7)
We start with an 8-hour burn-in on the GPUs and network, by running large GEMMs on the tensor cores and all-to-all communications at the same time, with a monitoring script tracking temperature, power, clock frequency, FLOPs, network connectivity, latency, bandwidth, and of course tailing the kernel ring buffer for any errors. You would be shocked how many hardware issues we induce with just this simple test. We cannot emphasize strongly enough how important it is to stress the GPUs and the network at the same time. Burn-in requires simultaneous thermal expansion and contraction of both of these components to approximate the real behaviour of these systems under real stress from real workloads. Lots of providers still use scripts that burn the GPUs and network separately. (2/7)
In our testing for ClusterMAX 3.0, one of the biggest things that we tested was reliability. we built on previous research to determine which providers can drive solid goodput by 1. identifying failures occurred, and 2. recovering from those failures. (1/7)🧵
The Chinese AI Infrastructure Boom:
Introducing the SemiAnalysis China Datacenter Model
1,000+ facilities across 60+ operators mapped,
built retail-first and flipped by AI,
largest hyperscaler leases 1/5 national capacity, 100MW in 12 months, Eastern Data Western Compute
https://t.co/twnBOg773d
MONEY PRINTER ALERT🚨 NVIDIA vLLM B200 CAN GENERATE UP TO💰️$15 BILLION💰️OF ANNUAL PROFITS PER GIGAWATT serving the open DeepSeekv4.1 Flash model at the official interactivity & official selling prices.
Using Engram DRAM offloading on NVIDIA results in a 50% increase in revenue per GigaWatt.