[Edaily Reporter Park Jung-Soo ] #Nota Inc., a company specializing in AI model lightweighting and optimization, has received consecutive recognitions for its optimization technology for ultra-large language models (LLMs) at a world-renowned artificial intelligence (AI) academic conference. Nota Inc. announced on the 25th that two papers on the optimization of ultra-large AI models were accepted for the main conference and the “Findings” section, respectively, at “EMNLP 2026,” an international academic conference in the field of natural language processing (NLP). Approximately 18,000 papers were submitted to this year’s EMNLP, with an acceptance rate of 15.4% for the main conference. EMNLP is an academic conference where global tech giants such as Google, Meta, and OpenAI, as well as researchers from major universities, present the latest research findings in natural language processing. The accepted papers addressed the performance degradation issues that arise during the quantization process of “Mixture of Experts (MoE)” models, which are currently being applied to the latest large language models (LLMs) such as the GPT family, Qwen, and Kimi. Although MoE is structured to compute only the necessary experts based on the input, the entire model must be stored in memory, requiring significant GPU and memory resources during actual operation. While quantization can reduce resource requirements, the model’s response performance may degrade if numerical changes cause a different expert to be selected than before. The paper accepted for the main conference proposed the “MENDS-MoE” methodology, which accounts for the impact of quantization on expert selection in subsequent stages and preserves the order of experts on the selection boundary. Experiments using 4-bit and 3-bit quantization on three types of MoE models showed higher average accuracy and language model performance than comparable techniques in most evaluation environments. A paper accepted to Findings proposed the “OPERA” methodology, which focuses optimization on changes that affect the final response rather than uniformly compensating for all expert selection changes. Both studies focused on minimizing performance degradation while maintaining appropriate expert selection even after quantization. Nota Inc. also placed third out of approximately 40 teams worldwide in the “Efficient Qwen Competition” held at the ICML 2026 “Adaptive Foundation Model Inference (AdaptFM)” workshop last July. By running the open-source LLM Qwen 3.5-4B on a single NVIDIA A10G GPU, the company maintained model performance while achieving an inference speed that was, on average, 6.978 times faster than previous methods. Two papers related to MoE quantization were also accepted at the same workshop. Nota Inc. is expanding the optimization technologies it developed during the research phase to ultra-large LLMs and data center AI infrastructure. As AI models grow in scale—with parameters ranging from hundreds of billions to trillions—technologies that reduce GPU and memory usage while maintaining performance are emerging as key factors determining the cost-effectiveness of AI services. Recently, the company optimized “Qwen 3.8 Max,” which has over 1 trillion parameters, reducing the number of NVIDIA B300 GPUs required to run it from 24 to 4. MoonShot AI also unveiled optimization results showing that its “Kimi K3” can reduce the number of B300 GPUs required from 8 to a maximum of 4, while Upstage’s “Solar Open 2” can reduce the number of H100 GPUs from 8 to 2. The company plans to expand its optimization capabilities—accumulated primarily in mobile and edge AI—into the domains of massive LLMs and data centers. To date, it has published over 50 research papers in major domestic and international academic conferences and journals, covering topics ranging from core technologies in pruning and quantization to mobile edge AI, generative AI, vision-language models (VLM), and LLMs. Chae Myeong-su, CEO of Nota Inc., stated, “As AI models grow to contain hundreds of billions or even trillions of parameters, the competitive focus of the AI industry is shifting from the performance of the models themselves to operational efficiency.” He added, “We will continue to invest in and support research and development to proactively respond to structural changes in the latest AI models, while expanding the scope of application for core optimization technologies that enhance the efficiency of AI infrastructure deployment.”
A union vote on the tentative agreement reached by SK hynix management and labor resulted in the agreement being rejected, with 50.1% voting against it.According to industry sources on the 25th, the e…
RN2 Technologies Co., Ltd.(148250)announced on the 25th that it has completed the development of a Proof of Concept (POC) for a global payment system for stablecoin-based tourism vouchers in collabora…
Creative works in which teenagers explored the proper use of AI and digital technology have been gathered in one place.LG HelloVision(037560)The Viewers Media Foundation announced on the 25th that it …