AI Weekly Malaysia

Summaries

Short AI and tech summaries with source links, signal scores, and why each update matters for builders, founders, and Malaysian tech workers.

Reset

Showing 1-3 of 3 results

DateProviderScoreSummary
01 Oct 2026, 7:30 PMTom's Hardware6.5 AI agents inadvertently leak 13,000+ internal screenshots from organizations

Tom's Hardware reports that AI agents inadvertently leaked more than 13,000 internal screenshots belonging to 300 organizations, with the exposed list said to include Fortune 500 companies and a frontier AI lab. The item was published 2026-10-01, but the text supplied here is almost entirely Tom's Hardware navigation and subscription markup — no leak mechanism, vendor name, storage location, discovery method, or response timeline is included.

Why: The only hard facts available are the counts (13,000+ screenshots, 300 organizations) and the claim that a frontier AI lab and Fortune 500 firms are on the list; the excerpt does not say which agent product, which storage path, or how the screenshots became reachable. So you cannot yet map this to your own stack — the defensible action is narrower: inventory which of your agents capture screenshots or browser state, and check where those captures are written and who can read them. Anyone running computer-use or browser-automation agents should treat screen captures as a data-exfiltration surface, not as throwaway debug output.

01 Oct 2026, 6:23 AMLatent Space6.5 Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week

Latent Space's DevDay 2026 episode (its first DevDay pod) features OpenAI's Computer Use (CUA) team and API platform leads, pushing back on Dwarkesh Patel's June 27, 2026 argument that computer use progress has been slow because the domain is 'clearly verifiable.' The counter-frame offered is that 'grindability is just as important as verifiability,' with guest Ari Weinstein (Sky cofounder, now working on CUA) describing computer use as '180 degrees different' from months ago as agents learn to debug and recover from failures. Concrete shipped artifact referenced: Computer History in ChatGPT, released Aug 14, 2026, which lets ChatGPT learn from everything you do on your computer, with a timeline view for reviewing that history.

Why: Two decisions here. First, the episode's stated architecture claim is that combining screenshots with accessibility data, the DOM, Playwright, and generated code is what changed the speed of computer-use agents - if you're building or evaluating an agent that drives a browser, that's a direct input into how you wire it up, versus screenshot-only loops. Second, Computer History (Aug 14, 2026) makes reviewable screen-activity capture a shipped consumer default in ChatGPT, so if you ship anything that records user screen or workflow data, users will now compare your privacy controls against a timeline view they can inspect. No Malaysia or SEA angle appears in the text.

28 Sep 2026, 5:44 PMHugging Face Blog6.5 Holo4: powering generalist computer-use agents

H Company released Holo4, a family of computer-use agent models in two sizes — 27B dense and 35B-A3B Mixture of Experts — plus Holotron4 Nano, an updated Holotron 3, all served on the H Models API with FP16, FP8 and GGUF weights on Hugging Face. The same model drives GUIs, writes and runs its own code, and calls MCP or API tools rather than needing a separate model per interface, trained via supervised and reinforcement learning on environments including ones generated by their Agentic Task Factory. On OSWorld 2.0 the 27B scores 61.7% against 81.8% for Opus 5.5, while the larger 35B-A3B MoE reaches only 30.9%, and every trajectory behind the published scores is open-sourced for replay or download.

Why: The open weights plus open trajectory dataset mean you can self-host a computer-use agent or fine-tune on their published steps instead of paying frontier API rates — but the size naming is a trap: the 35B-A3B MoE scores 30.9% on OSWorld 2.0 versus 61.7% for the 27B dense, so defaulting to the 'bigger' model for GUI work costs you roughly half the success rate. Pick the 27B dense or Holotron4 Nano for screen-based tasks, and check the FP8/GGUF builds against your own workflow before committing.

Top