AI Weekly Malaysia

Summaries

Short AI and tech summaries with source links, signal scores, and why each update matters for builders, founders, and Malaysian tech workers.

Reset

Showing 1-2 of 2 results

DateProviderScoreSummary
07 Oct 2026, 9:21 PMHugging Face Blog4.5 Introducing Falcon ASR

TII (Abu Dhabi) released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic with a stated emphasis on the Emirati dialect, plus English, French, Spanish and Portuguese. It reports 20.92% average WER and 8.79% CER across six Arabic test sets on the Open Universal Arabic ASR Leaderboard protocol, versus the best published snapshot figure of 23.17% (Audar-ASR-V1-Turbo, 2.35B params), and says it recorded the lowest word and character error rates among compared systems on an internal Emirati evaluation. It also outputs word-level timestamps and is available via a Hugging Face demo.

Why: The headline number is a 2.25-point WER improvement over the previous best leaderboard entry, but 20.92% still means roughly one in five words wrong on average Arabic test sets — fine for search or summarisation over audio, risky for verbatim transcripts, subtitles, or anything a user will read word-for-word. Competitor figures come from a 30 September 2026 leaderboard snapshot and the Emirati comparison is TII's own internal eval with no competitor numbers published, so if you are picking an ASR vendor for Arabic you should re-run your own audio rather than trust the table. The 1.6B size and open weights are the practical part: it is small enough to self-host next to your app if you would otherwise pay per-minute cloud ASR for Arabic, English, French, Spanish or Portuguese.

06 Oct 2026, 2:44 PMHugging Face Blog4.0 Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance

Falcon-Emirati-7B is a dialect-specialized LLM built on Falcon-H1-Arabic, targeting Emirati Arabic vocabulary, tone, and cultural context such as nabati poetry and proverbs. The post describes the underlying Falcon-H1 hybrid architecture, which runs Mamba and Transformer attention in parallel inside every block, and the broader family at 3B/7B/34B parameters with context windows up to 128K and 256K tokens. It is a vendor article with 9 upvotes and no benchmark results or independent evaluation in the excerpt.

Why: No direct Malaysia angle, but for Malaysian builders targeting Arabic-speaking or Gulf users, this is a 7B candidate to test on Emirati dialect prompts where Modern Standard Arabic misses meaning; since the post has no benchmark scores, run your own side-by-side eval before adopting it over a general Arabic model.

Top