cancel
Showing results for 
Search instead for 
Did you mean: 
angelborroy
Community Manager Community Manager
Community Manager

On 7 October 2026 I was in Belval for the Open Source Conference Luxembourg 2026. The theme was "Leading by example in Digital Sovereignty". The program had 89 talks on topics such as public sector, governance, security, decentralisation, EdTech and AI.

First, a big thank you to the organizers for having us. The venue, the University of Luxembourg campus in Belval, was a great place for a conference, and it is easy to reach by train. The organization was excellent from start to end. Many speakers also shared their slides on the pretalx schedule.

I spent most of the day in the EdTech and AI tracks. This post covers what I saw there and why it matters for Alfresco developers and for the Hyland OpenArch community.

conference.png

 

The theme of the day: LLMs on-prem, in Europe

One topic came up in almost every room. Everyone talked about running LLMs on their own infrastructure, in Europe. This was not a side topic. It was the main question for universities, ministries, media companies and software vendors.

The reasons were the same in every talk:

  • Data cannot leave the organization, or the country.
  • Licenses and terms of service can change at any time.
  • The cost of hyperscaler AI is hard to predict.
  • Regulation (GDPR, the EU AI Act) asks for control and evidence.

The question was not "can we run a model on-prem?" anymore. It was "which open model, on which hardware, and how do we operate it?".

EdTech track

Data That Travels, Larry Kilroy (DataKind). DataKind builds open source data infrastructure for more than 100 institutions that serve over one million students. The stack uses Apache Spark, Airflow and ECharts. One product, Edvise, predicts which students are at risk and explains why. Their main point: proprietary data models create vendor lock-in. A shared, open data model works as a "Rosetta Stone" for future AI tools. In early tests, reporting time went from 40 days to 4 hours.

Apps: a free software platform for the French Ministry of Education, Benoit Piedallu. After the 2020 lockdown, the ministry built a set of collaboration tools for 1.2 million staff. It runs entirely on free software. The ministry also contributes back to the upstream projects. In the Q&A, I asked which LLM they use. The answer: each service uses its own model today. They will likely need something to centralize this in the future. I think many organizations are in the same place.

Leaving the hyperscalers without leaving the game, Thibault Milan (Paperjam). This talk was a "logbook" of a real decision: should a B2B media company build its AI products on open models? Thibault shared a decision framework. It covered sovereignty, cost, control, model maturity, and the skills you need in-house. The message was honest. Open source AI is a strategic choice, not an ideology.

One talk I missed but recommend: [Ai]lice, by Nathan Lenas (LMDDC). It is a sovereign RAG (retrieval-augmented generation) chatbot for teachers, licensed under AGPL. Teachers upload documents or link a Moodle course, and every answer cites its sources. It has run in production since December 2025 on infrastructure run by LMDDC. This is very close to the use cases that Alfresco customers in education ask about.

AI track

ANN, RAG, MCP, or Agentic AI: It's All About Data, Marc Linster. Marc followed the path from plain LLMs to vector search, to MCP (Model Context Protocol), to agentic loops. His key line: models change every six months, but your data infrastructure stays. He grouped models into hosted APIs, open weights, and "true open source", and Olmo was in the last group.

Building a Sovereign, Open-Source Data and AI Stack for Europe, Dimitrios Avramidis (AIONIO). A full data and AI stack on EU clouds: DuckDB, SQLMesh, OpenTofu and Open WebUI with hybrid RAG.

Not to miss: The "Freeware" Fallacy of AI, Alfonso Cancellara (Red Hat). For me, this was the most interesting talk of the day. Alfonso compared popular models with three openness frameworks: the Stanford Foundation Model Transparency Index (FMTI), the Linux Foundation Model Openness Framework (MOF), and the OSI Open Source AI Definition. In his table, GPT-OSS, Gemma, Llama, Qwen, DeepSeek and Mistral Voxtral all failed the OSI definition. Only Olmo passed. His line stayed with me: "Open weights is to open source AI what freeware is to open source." The session was not recorded, but the slides are online. Read them before you choose your next model.

Your App Just Got a Second User, Sanjay Krishna Anbalagan. Treat the AI agent as a second user of your app. The app itself still decides what the agent is allowed to do. This is the same idea as my own talk (more on it below), applied to agents.

Beyond my tracks

Two other parts of the program are worth a mention. The day opened with Minister Stephanie Obertin, then The EU Open Source Strategy by Gemma Carolillo and a voluntary open source assessment framework by Marco Minghini. For a mid-size, community-driven conference, that is a strong signal: in Luxembourg and in Europe, open source is now public policy.

The agenda also had The Cyber Resilience Act and Open Source, with the CRA conformity journey of LibreOffice as a working example. It was at the same time as my talk, so I missed it. The Cyber Resilience Act (CRA) is the new EU law on security requirements for software products. It will matter for every open source project, Hyland OpenArch included, so I will read up on it.

My session: Sovereign RAG on your own repository

My talk was Sovereign RAG on Your Own Repository: Letting the Search Engine Enforce Permissions. If you point a RAG assistant at a document repository, it will quote documents that the user is not allowed to open. So the talk asked one question: who runs the permission check?

I walked through the history: the Alfresco Solr ACL join, then READER and DENIED fields in Elasticsearch, then OpenSearch Document Level Security, and finally Hyland OpenArch. In Hyland OpenArch, the search service applies the permission filter on the server, across Alfresco and Nuxeo. Everything ran locally: embeddings, search and the LLM.

Three patterns to take home:

  1. Put authorization in the storage layer, not in the application.
  2. Fuse ranks, not scores, with Reciprocal Rank Fusion (RRF), and measure on your own corpus.
  3. Use two-phase ingestion that is safe to re-run.

The slides and the handout are available. The code is in content-lake-app and content-lake-app-deployment. There is also a short demo video of OpenSearch Document Level Security. Content Lake is the open source proof of concept built on Hyland OpenArch: it adds the connectors, the ingestion pipeline and the RAG service on top of the Hyland OpenArch search service.

My takeaway: moving Content Lake to Olmo

My demo ran on Qwen 2.5. After the talks by Alfonso and Marc, I decided to change this. Content Lake now uses Olmo 3 7B Instruct from Ai2 (Allen Institute for AI) as the default model.

Why? Content Lake keeps documents, index and inference on infrastructure that you run. The model should follow the same rule. Olmo is open in full, not only open weights. Ai2 publishes:

  • The weights, under Apache 2.0
  • The pre-training data (Dolma 3) and post-training data (Dolci)
  • The training code (OLMo-core)
  • Checkpoints of every training stage
  • A technical report

You can see what the model was trained on, and no vendor license can change under you.

It also runs where we need it. On a laptop, the 7B model runs with Docker Model Runner as a 4.2 GB quantized file. On a GPU host, it runs with vLLM on a single NVIDIA A10G GPU. The setup is in content-lake-app-deployment.

What this means for Alfresco developers and Hyland OpenArch

  • Permissions are the hard part of enterprise RAG. Alfresco developers already know ACLs (access control lists) well. That knowledge is now key for AI. Do not leave the permission check to the application or to the LLM.
  • On-prem AI is expected. Many Alfresco customers in Europe are in the public sector, education and healthcare. They will ask for local models. The tools are ready: Docker Model Runner, vLLM and open models like Olmo. The next step is to share them. When every service brings its own model, you get duplicate GPUs, different answers and no single place to apply permissions.
  • Check how open a model really is. An open license does not make a model open source. Use FMTI, MOF or the OSI definition when you choose.
  • Education is a big content use case. Moodle, national platforms and teacher chatbots all need document management with permissions.
  • Open source needs a community. DataKind invited institutions to join their community, and the French ministry contributes back to the projects it uses. Open source projects grow when people build them together.
This is exactly what Hyland OpenArch provides. It answers the questions I heard all day in one open source platform: complete on-prem deployment, permissions enforced by the search service, multi-source search with one shared index for Alfresco, Nuxeo and other sources, and your choice of LLM, from Olmo on a laptop to vLLM on a GPU server. It is also the data layer that stays when models change: you can swap the LLM without re-indexing your content. And AI agents get the same permission-filtered results as the user they act for.

This is where the Hyland OpenArch community comes in. Hyland OpenArch is open source under Apache 2.0. Content Lake already has connectors for CMIS and the filesystem. Is your source missing? Create your own connector: content-lake-app includes a guide, a Maven archetype that generates the project skeleton, and a worked example. We want to build it with you: try it, open issues, share your use cases and send pull requests.

The real value: conversations

The talks were good, but the best part of the day happened between them. I met a developer working on hestIA, an on-prem, multi-tenant RAG assistant for policy and compliance documents. We are working on the same problem from two sides, so we shared implementation details and looked for ways to collaborate.

The two projects fit together well. This is how Hyland OpenArch differs from the other approaches I saw during the day:

  1. It is a layer, not a chat app. Assistants like hestIA or [Ai]lice are the apps that people use every day. Hyland OpenArch is the content and permissions layer below them, so an assistant can run on top of it.
  2. Permissions come from the source, per document. Other approaches control access per tenant, per course or per collection, or leave the filter to the application. Hyland OpenArch keeps the real ACL of each document from Alfresco or Nuxeo and enforces it in the search service.
  3. Connectors, not uploads. Uploaded copies drift from the source. Content Lake connectors sync from the live repository into Hyland OpenArch, so when a permission changes, the index follows.

I also met a colleague from Passbolt, an open source password manager for teams, made in Europe (source code on GitHub). He brought the best stickers I have ever seen at a conference. I am going to test Passbolt with my students at the University.

Best stickers ever!Best stickers ever!

 This is the real value of a mid-size, community-driven conference like this one. You meet the people who are building the same things as you, and you have time to talk.

Thanks, Luxembourg

Thanks again to the organizers, the speakers and everyone who contribute to make this a great event. It was a great day of talks and conversations. See you next year in Belval.