Францын H стартап график интерфэйсийг удирдах чадвартай хиймэл оюун ухааны шинэ загваруудаа танилцууллаа

Published:

Энэхүү мэдээ, нийтлэлийг хиймэл оюун боловсруулав.

Энэхүү загварууд нь Windows болон Linux үйлдлийн систем дэх график орчинд хулганаар заах, товших, гүйлгэх зэрэг үйлдлийг бие даан гүйцэтгэх боломжтой юм

Францын хиймэл оюун ухааны H компани даваа гарагт график хэрэглэгчийн интерфэйс (GUI) дээр ажиллах зориулалттай Holo 4 загварын гэр бүлийг танилцууллаа. Өмнө нь хиймэл оюун ухааны агентууд ихэвчлэн тушаалын мөр (CLI) болон API-аар дамжуулан ажилладаг байсан бол шинэ загвар нь хэрэглэгчийн үйлдлийн систем дэх дүрслэл бүхий орчинд илүү уян хатан ажиллах боломжийг олгож байна.

Holo 4 загваруудыг Alibaba компанийн Qwen 3.8 27B болон Qwen 3.6 35B-A3B загварууд дээр суурилан, хяналттай сургалт болон бататгах сургалтын аргаар бүтээжээ. Мөн тус компани NVIDIA-ийн Nemotron 3-т суурилсан Holotron загвараа шинэчилж, GUI удирдах чадвараар сайжруулсан байна. Туршилтын явцад Holo 4 нь FreeCAD программ хангамжийг ашиглан 3D загвар зохион бүтээх зэрэг нарийн даалгавруудыг гүйцэтгэжээ.

Тус компанийн мэдээлснээр, Holo 4 нь илүү том хэмжээтэй загваруудаас өндөр гүйцэтгэлтэй боловч ашиглалтын явцад илүү их тооцоолол шаарддаг тул нэг даалгаврын өртөг өндөр байж болзошгүй байна. Гэсэн хэдий ч загваруудын хэмжээ харьцангуй бага тул 24 GB RAM бүхий Nvidia RTX 3090 зэрэг дундаж түвшний техник хангамж дээр ажиллуулах боломжтой.

Хэрэглэгчид Hugging Face платформоос Holo 4-ийн GGUF хувилбарыг татаж авах боломжтой бөгөөд үүнийг Llama.cpp, LM Studio, Ollama зэрэг хэрэгслээр ашиглах боломжтой. Түүнчлэн H компани өөрийн нээлттэй эхийн HAI-Agents хэрэгслийг GitHub дээр байршуулсан бөгөөд ирээдүйд DSpark технологиор дамжуулан тооцооллын хурдыг нэмэгдүүлэхээр төлөвлөж байна.

Дэлгэрэнгүйг эх сурвалжаас харах

↓Эх сурвалжийг нээх ↓

Most LLMs are great at answering prompts, but fall short when it comes to navigating around the desktop in Windows or Linux. French AI model dev H unveiled a pair of computer use models on Monday aimed at handling graphical user interfaces (GUIs). Throughout computing history, computer use largely falls into three categories: command line interfaces (CLIs), application programming interfaces (APIs), and GUIs. AI agents can easily plug into the first two, but navigating desktop environments and applications that often prioritize form before function remains an ongoing challenge. H’s Holo 4 family of models aims to address this challenge by enabling relatively small but capable models to tackle all three computer use scenarios including pointing, clicking, scrolling, and typing their way through graphical interfaces originally meant for us meatbags. Fine tuned using supervised training and reinforcement learning, Holo 4 is built atop Alibaba’s Qwen 3.8 27B and Qwen 3.6 35B-A3B models. And by optimizing for CLIs, APIs, and GUIs, H claims that its models achieve far greater versatility than pure computer use models might otherwise. Alongside Holo 4, H has also updated its Holotron model, which is based on Nvidia’s Nemotron 3, with similar capabilities. In one example, the company showed Holo 4 27B taking advantage of FreeCAD’s macro function to programmatically design a 3D model of the Eiffel Tower rather than manually building it using primitives like cubes. In another demo, H did the opposite using extruded shapes to recreate the company’s logo, showing the model’s flexibility. As with any model dev’s benchmarks, take these claims with a grain of salt, but if H is to be believed, Holo 4 outperforms significantly larger frontier models from the likes of OpenAI, while using a fraction of the parameters. Curiously, this doesn’t mean that they’re cheaper. In fact, while the company shows higher scores, in many cases the models end up costing more per task. Given what we know about Qwen 3.8 27B, this is likely due to Holo 4 using substantially more “thinking” tokens in order to arrive at a final result relative to something like GPT 6 Luna, which doesn’t perform as well in the OSWorld 2.0 benchmark, but costs substantially less. Having said that, the open weights models’ diminutive size means that researchers, AI enthusiasts, and enterprises should be able to run them on relatively modest hardware. A 24 GB Nvidia RTX 3090 should be more than capable of running these models at 4-bit precision. H clearly expects users to do just that since alongside BF16, FP8, and NVFP4 weights, it’s also made a Llama.cpp (and by extension LM Studio and Ollama)-friendly GGUF version of the model available for download on Hugging Face. The model dev says that it also plans to release DSpark draft weights in order to speed up inference using a technique called speculative decoding. We’ve explored this performance-enhancing inference tech in the past, but in a nutshell it uses a small model to guess the outputs of a larger model. When it works, users experience a speedup in token processing and generation and, when it doesn’t, it falls back to the base model ensuring no loss in output quality. However, the models aren’t worth much without a harness. H has developed several agentic harnesses including its open source HAI-Agents harness, which is available for download on its GitHub. However, in theory the models should work with third-party computer use harnesses. H isn’t the only model dev focused on computer use applications. At AWS’ Re:Invent conference last year, the company announced its own set of computer-use models. Meanwhile, the big three American model labs, OpenAI, Google, and Anthropic, are also investing in this capability, perhaps because escaping their sandbox sometimes requires pushing a button. ®

Та юу гэж бодож байна?

Сэтгэгдлээ оруулна уу!
Please enter your name here

MFC.mn сайтад сэтгэгдэл оруулахад анхаарах зүйлс

Холбоотой

spot_img

Шинэ

spot_img