AI Weekly Malaysia

Summaries

Short AI and tech summaries with source links, signal scores, and why each update matters for builders, founders, and Malaysian tech workers.

Reset

Showing 1-15 of 15 results

DateProviderScoreSummary
30 Sep 2026, 5:55 PMHacker News8.0 You Said No MCP

Earendil Engineering published a post explaining why Pi reversed its public position on MCP: Pi's site and podcasts had previously declared that Pi does not support MCP, and MCP was available only as an extension, but it is now part of Pi's core. The post says the change came from rethinking MCP rather than from MCP improving alone, and that the same sandbox/interpreter work needed for MCP also makes it easier to use Jev inside Pi. Earendil argues MCP's biggest remaining problem is composition — even with codemode — and that MCP should be closer to OpenAPI with intelligent tool discovery, returning structured data instead of text.

Why: If you maintain an MCP server that returns prose text to save tokens, this post is a direct argument that you are optimizing for the wrong harness: Earendil says tools should return structured data and be discoverable by their documentation and description. It also matters if you build on Pi specifically, because MCP moved from optional extension to core, so upgrade behaviour changes rather than being opt-in. The composition complaint is the practical warning — even a core MCP implementation with a sandbox does not fully solve chaining tool calls, so plan for that gap rather than assuming the integration removes it.

29 Sep 2026, 2:08 PMThe Hacker News8.0 Official MCP Python SDK Flaw Can Let Malicious Servers Steal OAuth Credentials

The maintainers of the official MCP Python SDK published a security advisory saying a malicious MCP server could point a client at an attacker-controlled token endpoint, causing the SDK to send the client secret, authorization code, and PKCE proof key to the attacker instead of the real login service. Cycode, which reported the flaw, demonstrated the full exchange in a test and says the resulting access token carries whatever permissions the app was granted; because the client secret is long-lived, it keeps working until changed. Fixed in SDK versions 1.30.0 and 2.2.0; scored 7.5 for the two providers that run without a person present and 6.5 for the interactive provider, with no CVE assigned as of September 29.

Why: If your Python MCP client connects over HTTP using OAuthClientProvider or ClientCredentialsOAuthProvider on a version below 1.30.0/2.2.0, the server you connect to could have redirected your client secret, auth code, and PKCE key to itself — so upgrade, and then rotate the client secret, because the fix does not invalidate a secret that already leaked. Note the interactive case still requires a human to approve a page that Cycode says is the genuine login page, so user approval is not a defence here.

04 Oct 2026, 6:56 AMHugging Face Blog7.5 The Agent Said It Was Done. The Database Disagreed.

Microsoft and Hugging Face published ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than on their final sentences or tool-call validity, across 507 stateful business workflows each run 20 times per model. The illustrative retail case runs nine well-formed tool calls but fails a single executable check: the ticket's status is 'solved' where the required end state is 'hold'. The benchmark is available through Hugging Face and runnable via OpenEnv, with the specific task published as sandbox_external_retail_group1.py:test_case_ST003_006 and the full trace in Appendix D.4, Case 3.

Why: If your agent writes to tickets, orders, or account records, an eval that checks the reply text or that tool calls were well-formed will pass this exact failure: nine valid calls, wrong persisted value. The concrete fix here is asserting on the field the workflow must end in (this case: ticket status 'hold', not 'solved') and rerunning the same task 20 times, because the benchmark's whole premise is that one passing run says nothing about reliability. There is no Malaysia-specific angle in this text.

02 Oct 2026, 3:33 AMHacker News7.5 Pi 1.0

Earendil shipped Pi 1.0, a self-described 'hardened, minimal, extensible agent harness,' adding Codemode (native MCP support plus non-LLM models like Jev and image models), virtual-model extensions, deferred tool loading, cache warming for Anthropic models, mid-conversation system messages, a new TUI theme, and full-screen mode by default. Alongside it, Earendil released Pi Durable, an experimental package for building long-running agentic applications that it says shares Pi's minimalism but targets longer conversations and tasks outside the terminal. The post claims hundreds of thousands of weekly Pi users and says features were only adopted after months of being 'thrown up against the wall,' with a longer list of rejected ideas. The HN thread drew 1108 points and 344 comments.

Why: If you run coding agents against Anthropic models, cache warming and deferred tool loading are the two items here that touch your token spend and startup latency directly, and mid-conversation system messages change what you can do mid-transcript when swapping prompts or tool sets in a long session. Codemode adding native MCP support means an existing MCP server setup may work without a separate bridge, which is worth testing before writing glue code. Pi Durable is explicitly labelled experimental, so treat it as something to prototype on rather than to put a production long-running workload behind. Nothing in the text states any Malaysia- or Southeast Asia-specific pricing, hosting, or policy angle.

29 Sep 2026, 3:33 AMTechCrunch7.5 Shopify opens checkout to browser-based AI agents

Shopify announced that browser-based AI agents can now complete purchases on eligible merchants' sites, extending its earlier WebMCP support from product search and add-to-cart into checkout, including Shop Pay. The update ships three new tools — get_checkout, update_checkout, and complete_checkout — letting an agent read the checkout screen, change details like address or delivery option, and submit the order once the buyer authorizes it, without screenshots or page scraping. Shopify already runs a hosted MCP server for server-to-server agents; both paths sit on its Universal Commerce Protocol (UCP), and the feature is rolling out to all eligible merchants according to Shopify's Gil Greenberg, who works on agentic commerce. The move runs opposite to Amazon and Adidas, which the article says are blocking AI agents from purchasing on users' behalf.

Why: If you run a Shopify storefront, agent traffic can now finish checkout instead of stalling at the cart — so you have to decide whether to leave it enabled or block agents the way Amazon does, and check whether your checkout customizations survive an agent editing address and delivery fields. If you build agents or commerce tooling, this is a concrete interface to target: implement get_checkout, update_checkout and complete_checkout (or the hosted MCP server path) rather than driving checkout with screenshots and scraping. For founders evaluating agentic commerce in Malaysia or SEA, the practical question is whether agent-completed orders change your payment, fraud or fulfilment assumptions before you enable it.

02 Oct 2026, 2:40 PMLatent Space6.8 [AINews] Pi 1.0, Pi Durable, and AIE NYC

Earendil's Pi 1.0 and Pi Durable both landed on the Hacker News front page. Pi 1.0 ships Codemode (native MCP plus Jev and image model support), deferred tool loading, cache warming for Anthropic models, mid-conversation system messages that change prompts and tools mid-transcript, and full-screen TUI by default. Pi Durable is a TypeScript port that checkpoints every agent step so runs auto-resume after a crash, runs anywhere with a JS runtime (Node, Bun, Cloudflare) with Memory, SQLite, or JSONL storage, supports parallel branching conversations, installable Extensions with rollback-able durable tasks, background compaction, state documents shared alongside transcripts for multi-user steering, and hot-swapping tool code while the agent is running.

Why: The checkpoint-per-step plus pluggable storage model is a concrete design you can copy if your agents currently die with the process: state that survives a restart, and tool code you can hot-swap without draining a run, changes how you'd structure long multi-step workflows like a checkout flow with rollback. The shared state document next to the transcript is the detail worth stealing if you want more than one user or UI to watch and steer the same agent. There is no Malaysia-specific angle in this item, and the Gemini 4 Argon / GPT-6.1 Sol / FLUX 3 section is explicitly flagged as developer accounts rather than independent evidence, so treat it as unverified.

30 Sep 2026, 9:00 PMCloudflare Blog6.5 Monetization Gateway beta: charge AI agents for consumption with HTTP 402

Cloudflare opened a closed beta of its Monetization Gateway, which lets domain owners charge AI agents per use for websites, APIs, MCP tools, or datasets, built on the HTTP 402 payment-required pattern and the x402 flow. Cloudflare says it has worked with customers since announcing the plan three months ago and is showcasing four customer use cases it claims are in production today. Pricing is per request, per search query, or per token, and the post states stablecoin transactions on blockchain networks are what currently support those economics; access is requested through the Cloudflare Dashboard.

Why: If you sell an API, MCP tool, or dataset behind Cloudflare, this is a different billing model from subscriptions and prepaid credits: per-request/per-query/per-token metering with payment attached to the HTTP request itself. The concrete decision is whether your billing stack can handle sub-dollar, per-call charges and whether you can accept the stablecoin rail Cloudflare names as the current enabler — if not, the beta is not usable for you yet. Since it is a closed beta with no published pricing or fee schedule in the post, request Dashboard access only if agent traffic is already a real share of your usage.

28 Sep 2026, 9:00 PMCloudflare Blog6.5 Introducing Forge: the open source pipeline for generating SDKs, CLIs, docs, and more

Cloudflare open sourced Forge, a pluggable generation pipeline that produces SDKs, CLIs, docs, and libraries, free to deploy and run yourself. It already generates the output behind the cf CLI and is slated to power Cloudflare's API docs and SDKs over the next few months, built to handle an API with over 3,500 operations across services written in Rust, Go, TypeScript, and Python. Forge runs in CI on each team's API repos: it lints every change and produces a preview build of the CLI, docs, and SDKs with only that change highlighted, which developers can install and test before merging — the same premise as Workers Previews.

Why: If you maintain an API, the concrete takeaway is the per-PR preview model: instead of discovering a broken generator at release time, every API change gets an installable CLI/SDK/docs preview before merge. The post also states Cloudflare tried several hosted generation products and 'some have shut down entirely', so if you currently depend on a hosted codegen vendor for your SDKs or docs, that is a signal to check whether your pipeline is portable to something you run in CI. For agent-facing surfaces specifically, the stated motivation is that CLIs, SDKs, MCP servers, and docs are now expected for every product, not just developer products.

28 Sep 2026, 5:44 PMHugging Face Blog6.5 Holo4: powering generalist computer-use agents

H Company released Holo4, a family of computer-use agent models in two sizes — 27B dense and 35B-A3B Mixture of Experts — plus Holotron4 Nano, an updated Holotron 3, all served on the H Models API with FP16, FP8 and GGUF weights on Hugging Face. The same model drives GUIs, writes and runs its own code, and calls MCP or API tools rather than needing a separate model per interface, trained via supervised and reinforcement learning on environments including ones generated by their Agentic Task Factory. On OSWorld 2.0 the 27B scores 61.7% against 81.8% for Opus 5.5, while the larger 35B-A3B MoE reaches only 30.9%, and every trajectory behind the published scores is open-sourced for replay or download.

Why: The open weights plus open trajectory dataset mean you can self-host a computer-use agent or fine-tune on their published steps instead of paying frontier API rates — but the size naming is a trap: the 35B-A3B MoE scores 30.9% on OSWorld 2.0 versus 61.7% for the 27B dense, so defaulting to the 'bigger' model for GUI work costs you roughly half the success rate. Pick the 27B dense or Holotron4 Nano for screen-based tasks, and check the FP8/GGUF builds against your own workflow before committing.

30 Sep 2026, 9:00 PMCloudflare Blog6.0 Simplifying domains for people and agents

Cloudflare Registrar shipped a redesigned domain search that lists all 420+ supported extensions with live-as-you-type results, sorting, filtering, and transparent at-cost pricing, plus an expanded Registrar API, MCP integration, and a new cf CLI. The API, in beta since April, now adds a sandbox that tests search/check/register workflows without a real transaction or purchase, an extensions endpoint covering registry-specific requirements across those 420+ TLDs, and programmatic transfer-in with an EPP auth code. Cloudflare says you can now prompt an agent to run commands like `cf registrar registrations check example.com`, `create`, or `transfer-in ... --auth-code`.

Why: This is the first mainstream registrar where the buying flow is designed to be driven by an agent or script rather than a checkout page, so if you register domains manually today you can decide whether to move that step into your agent/tooling stack. Two details change what's safe to do: the sandbox means you can build and test the full register-and-transfer path without spending money, and the extensions endpoint is what you need to handle per-registry requirements instead of hardcoding .com assumptions. Note that transfers are now scriptable, which makes it practical to move existing domains off an upsell-heavy registrar in bulk.

29 Sep 2026, 9:07 PMHugging Face Blog6.0 Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

A Hugging Face blog post from MultiverseComputingCAI (Antonio Tiene, Ander Alvarez Sanz, Oliver Wirjadi) introduces ProvenanceGuard, a factuality verifier for MCP-based LLM agents that checks not just whether a claim is supported by pooled evidence but whether the supporting source matches the source the answer names. It targets a failure mode the authors call 'cross-source conflation' — e.g. a 30-day refund window that is real but stated in a policy document while the answer attributes it to the account record, or a patient-history detail presented as a medical-literature finding. The post argues existing checkers (RAGAS faithfulness, MiniCheck, AlignScore, SummaC) pool evidence and therefore pass such claims, and points to a paper on Hugging Face and arXiv, though the excerpt cuts off before any accuracy numbers or benchmarks.

Why: If you ship an MCP agent that writes citations like 'according to the account record', RAGAS-style faithfulness scoring will not catch a claim that is true in some other tool output but attributed to the wrong one — and in support, clinical, or financial contexts that misattribution is as damaging as a wrong fact. The practical decision is to add a per-source check (does the cited tool output actually contain the claim?) rather than a pooled-evidence score; note the post publishes no measured improvement over the existing checkers, so treat it as a design pattern to prototype, not a drop-in library to adopt.

29 Sep 2026, 8:30 PMTechCrunch6.0 Reco raises $55M as AI agent security startups crowd the market

Reco raised $55M and repositioned from mapping/ securing SaaS and AI platforms to a broader 'context graph' product that links AI agents to apps, people, accounts and permissions so security teams can see what an agent can reach and revoke unnecessary access. TechCrunch notes at least two dozen companies now sell some form of AI agent security, with vendors converging on similar pitches (knowledge graphs, continuous monitoring, runtime security, MCP vetting), including CrowdStrike building detection and response on the devices agents run on. Reco's CEO Ofer Klein says its platform found 21,000 unknown agents at one Fortune 100 customer, and at a large financial services customer it found an agent created by an ex-employee that could access Salesforce and share data to an unseen domain.

Why: The concrete number to act on is the 21,000 unknown agents found at a single Fortune 100 company: if you have been shipping agents with service accounts, OAuth scopes or MCP tool access, you probably cannot enumerate them today, and an ex-employee-owned agent with live Salesforce access is the failure mode. That also means agent-inventory and permission-graph tooling is now a crowded category with two dozen-plus vendors, so if you are a founder eyeing this space, differentiate on a specific surface (MCP tool vetting, runtime revocation, data egress) rather than a generic 'discover and govern agents' pitch. Note the numbers come from Reco itself, not independent measurement.

28 Sep 2026, 9:00 PMCloudflare Blog6.0 EmDash 1.0: the stable CMS with a secure plugin registry

Cloudflare shipped EmDash 1.0, a free, open-source CMS built on Astro, after teasing it on April 1 as a "spiritual successor to WordPress" and spending roughly five months hardening data safety, database migrations, editorial workflows, localization, plugin security, and the admin/API/MCP/media paths. Editors work in the EmDash admin, developers build with Astro, and agents operate through the API, CLI, or a built-in MCP server. It also launches a decentralized plugin registry, where developers publish plugins without giving up ownership of their identity or releases to a central marketplace and site owners install them from inside EmDash; the post cites Avulux moving a custom microsite off WordPress in under a day using EmDash Agent Skills. The article is truncated mid-sentence in the migration section, so no independent performance, security, or upgrade-compatibility data is provided.

Why: If you maintain WordPress sites for clients or sell site-building as a service, EmDash 1.0 is now a free Astro-based option with an MCP server and CLI that agents can drive — meaning the editing interface is no longer only a human admin panel. The concrete trade-off is the plugin registry: with no central marketplace, there is also no central review, signing authority, or takedown process, so vetting plugin provenance becomes your responsibility before it touches a client site. The only migration evidence in the post is a vendor-cited customer (Avulux, under a day), so treat the speed claim as unverified and pilot on one non-critical site before committing a client's content.

28 Sep 2026, 9:00 PMCloudflare Blog6.0 The road to the agentic browser: A Kitesurf update

Cloudflare published a follow-up on Kitesurf, the browser it launched in August that runs entirely on Cloudflare Workers, and the headline change is WebMCP support: sites can expose callable tools (the post uses searchFlights() and Cloudflare Radar's navigate-to / set-location as examples) so agents call functions instead of simulating clicks. Kitesurf also added a batch of browser standards — CSS Layout, CSSOM, CSS Typed OM, custom elements, plus URL-based module resolution, JSON modules, and import map handling — aimed at rendering heavier JavaScript-chunked pages. You can try it in Cloudflare's public Kitesurf playground or point an agent at it via a chrome-devtools-mcp config using a wss:// browser-run devtools endpoint with browser=kitesurf and --category-experimental-webmcp.

Why: If you build or operate agents that touch websites, this gives you a concrete alternative to pixel-clicking: check the Application tab in DevTools on a target site to see whether it already publishes WebMCP tools, and if it does, wire your agent to call them. The MCP config in the post (chrome-devtools-mcp@latest pointed at the browser-run WebSocket endpoint with --category-experimental-webmcp) is copy-pasteable, so you can test tool-calling against Radar without running your own headless browser. Caveat: this is Cloudflare announcing its own product and the post claims Kitesurf is now 'more capable and more efficient' without publishing any benchmark numbers, so treat the efficiency claim as unverified.

30 Sep 2026, 1:15 AMTechCrunch5.0 OpenAI expands ChatGPT’s plug-ins with app-like interfaces and automations

At its Dev Day on September 29, 2026, OpenAI said developers can now build app-like plugin extensions inside ChatGPT, giving apps a dedicated slot in the ChatGPT sidebar plus interactive panels and file viewers that support whatever formats the developer's product uses. It also introduced a Plugin Creator tool, a redesigned submission flow with clearer feedback, improved plugin ranking and recommendation in the directory and in conversations, per-plugin permission approval, plugin hosting on ChatGPT Sites, and support for the proposed MCP Events specification so plugins can trigger automations from events in a connected app.

Why: If you already maintain an MCP server or ChatGPT plugin, two changes are worth acting on: plugins can now live in the sidebar with their own interactive UI, and MCP Events support turns them from on-demand tool calls into event-driven automations. The text gives no pricing, availability dates, rate limits, or migration path for existing plugins, so treat this as something to prototype against rather than something to schedule a release around. There is no Malaysian or Southeast Asian angle in this article.

Top