BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CERN//INDICO//EN
BEGIN:VEVENT
SUMMARY:CLARIN.SI Open language data for AI development and evaluation in 
 Slovenian and other South Slavic languages
DTSTART:20261111T080000Z
DTEND:20261111T160000Z
DTSTAMP:20260928T002400Z
UID:indico-event-3826@indico.ijs.si
CONTACT:events@slaif.si
DESCRIPTION:Speakers: Taja Kuzman Pungeršek\, Nikola Ljubešić\n\nCourse
  provider: Jožef Stefan Institute (JSI)\, Department of Knowledge Technol
 ogies (E8)Participating organisations: Common Language Resources and Techn
 ology Infrastructure CLARIN.SI\; CLARIN knowledge centre for South Slavic 
 languages CLASSLAInstructors: Nikola Ljubešić (E8 JSI)\, Taja Kuzman Pun
 geršek (E8 JSI)\nLanguage data are the crucial foundation for the develop
 ment and evaluation of modern language technologies\, which include large 
 language models\, automatic speech recognition models\, machine translatio
 n systems and chatbots. Training data determine which languages a model ca
 n handle and what kinds of biases are reflected in its behavior. High-qual
 ity data are essential not only for initial model training\, but also for 
 adapting models to specific tasks\, such as question answering or text cla
 ssification\, where large volumes of carefully designed and manually annot
 ated examples are required. Finally\, reliable and representative data are
  indispensable for evaluation of the model capabilities in the target lang
 uage\, such as the Slovenian language. This lecture introduces open langua
 ge data provided through the CLARIN.SI Trust Core certified repository as 
 a key resource for responsible and effective development of language techn
 ologies.Learning objectives: The main objective of the course is to inform
  small enterprises and language technology developers about the availabili
 ty of open language data that can be used across all stages of AI developm
 ent and evaluation. Participants will become familiar with the CLARIN.SI i
 nfrastructure\, which serves as the central national hub for language reso
 urces in Slovenia.Course content: The lecture provides an overview of open
  language data types used in language technology development\, including m
 assive text corpora\, speech data and evaluation datasets. Special attenti
 on is given to the CLARIN.SI repository\, which contains approximately 700
  language resource entries\, around 400 of them focused on the Slovenian l
 anguage\, together comprising about 9 terabytes of data. The lecture prese
 nts concrete examples of widely used resources\, such the large web-based 
 text corpora used for training large language models\, instruction-tuning 
 and task-specific datasets\, automatic speech recognition resources\, and 
 evaluation benchmarks for Slovenian and other South Slavic languages.Learn
 ing outcomes: After the lecture\, participants will understand the central
  role of language data in shaping the performance and reliability of langu
 age technologies. They will gain a clear overview of what the CLARIN.SI re
 pository offers and how its resources can be effectively leveraged for the
  development and evaluation of AI systems for Slovenian and South Slavic l
 anguages. Participants will also be aware that the CLARIN.SI repository ca
 n be used to deposit their own language data\, ensuring long-term archivin
 g\, increased visibility and reuse\, and compliance with data management p
 lan requirements.\n\nhttps://indico.ijs.si/event/3826/
IMAGE;VALUE=URI:https://indico.ijs.si/event/3826/logo-1596601024.png
LOCATION:Slovenija
URL:https://indico.ijs.si/event/3826/
END:VEVENT
END:VCALENDAR
