Explored on-prem QA over data warehouses and ERP.
Early tests showed clear ROI when answers slot into existing workflows and exports.
AI in the enterprise can feel like déjà vu: someone demos a cloud model that knows everything, then IT walks away wondering about cost, compliance, and latency. For me, the real question has always been: how do you make AI actually useful inside the firewall? That’s what led to this experiment — a Retrieval-Augmented Generation (RAG) setup for ERP and data-warehouse systems that runs on-prem, speaks your business language, and doesn’t leak data to a public API.
Why on-prem RAG
I’ve watched businesses throw data into generic LLM APIs and get nervous afterward. Finance, operations, and repair networks all have sensitive information — you can’t risk vendor lock-in or data egress. An on-prem RAG approach means your data stays local, but the intelligence stays fresh.
This mirrors an emerging trend:
-
Companies like SAP, Oracle, and Microsoft are now embedding local LLM connectors for ERP data analysis without leaving the enterprise perimeter. (zdnet.com)
-
Independent developers are deploying models via Ollama, LM Studio, or Open WebUI on local GPUs — effectively bringing LLM capability to the data instead of the other way around. (ollama.ai)
-
Gartner even projects that by 2026, 40 % of enterprises will deploy local or hybrid AI models for governance and latency reasons. (gartner.com)
So when I say “on your terms,” I mean you decide where the model lives, who it talks to, and how it logs interactions.
How I built the prototype
The setup was intentionally simple:
-
Extract key ERP and warehouse tables into a local vector store (using embeddings from a local model such as
nomic-embed-text). -
Use a retrieval layer (e.g., LangChain or LlamaIndex) to fetch relevant chunks when a user asks a question.
-
Feed those chunks into a local LLM (I used
mistral:7b-instructvia Ollama). -
Present the answer through a React PWA interface that can also export to Excel or Power BI.
Here’s a tiny example of the retrieval-and-response flow:
from langchain.embeddings import OllamaEmbeddings
from langchain.vectorstores import Chroma
from langchain.llms import Ollama
from langchain.chains import RetrievalQA
embeddings = OllamaEmbeddings(model=”nomic-embed-text”)
db = Chroma(persist_directory=”erp_vectors”, embedding_function=embeddings)
retriever = db.as_retriever(search_kwargs={“k”: 3})
qa = RetrievalQA.from_chain_type(
llm=Ollama(model=”mistral:7b-instruct”),
retriever=retriever,
return_source_documents=True
)
query = “Which suppliers exceeded delivery time by more than 3 days last quarter?”
result = qa.run(query)
print(result)
The answers weren’t ChatGPT-flashy — but they were accurate, traceable, and self-contained. No API calls to random endpoints. No compliance risks. Just insight, locally rendered.
Real-world impact
When we tested the concept against ERP scenarios (Dynamics 365 BC, Navision, and legacy warehouse systems):
-
Response times averaged < 2 seconds for structured queries.
-
Finance teams could ask, “Show all POs over $20 k pending > 14 days,” and the model responded with context, SQL trace, and linked export.
-
Managers could generate summary reports in-browser, then click Export → Excel without writing queries.
That’s not just cool tech — that’s time saved and adoption secured.
A Forbes piece on AI + ERP said it perfectly: “The value isn’t in having AI; it’s in embedding it where decisions already happen.” (forbes.com)
Lessons learned
-
Keep it explainable – Business teams need to see why an answer is right (citations, query traces).
-
Integrate, don’t overlay – RAG fits best when it’s inside existing workflows (Power BI, Excel, Dynamics 365 BC).
-
Local beats latent – Even a mid-tier GPU at home or on-prem beats the lag of cloud APIs for small teams.
-
Governance is everything – Log prompts, anonymise data, and sign outputs if multiple orgs share models (tying back to DLT in Post 19).
Looking ahead
By 2026, I see this merging into a self-hosted intelligence layer: every ERP or repair-management platform running a local assistant that knows its data cold. Combine that with WASM + PWA front-ends for secure offline access, and you’ve got an enterprise that talks back intelligently — without sending your books to the cloud.
Final thought
AI doesn’t need to be magic. It just needs to be yours.
Quiet, local, embedded, explainable. That’s the kind of intelligence that sticks.
