Skip to content
FenixIAm by AIworks 2028 FenixIAm

Intelligence Centre September 2026 edition 6 pieces

What you need to know before putting AI to work.

Guides and analysis written from the experience of building agents for real businesses: how to measure, what the rules require, which model to choose and where AI is best left out. For the people who have to decide, without the hype.

01Training

How to score a sales role-play without fooling yourself

· 2 min read

A score without the sentence that backs it up does not help anyone improve. Observable criteria, literal quotes and progress over time: that is how a rehearsal becomes learning.

Criteria you can hear

The first mistake is scoring what cannot be seen. “Positive attitude” or “confidence” depend on who is listening. A good criterion describes a behaviour you can point to in the transcript: “asks about the need before presenting the product”, “rephrases the objection before answering it”, “proposes a next step with a date”. Eight to ten criteria are enough; beyond that, the report turns into noise.

Every criterion needs an anchored scale. If it runs from 0 to 10, you have to write down what a 3, a 6 and a 9 look like, with examples. Without anchors, two raters, human or not, will score the same conversation differently.

Every score, with its quote

A score is only useful when it comes with the sentence that justifies it and the minute it was said. The quote lets the salesperson understand what went well or badly, and lets the trainer check whether the assessment is fair. If the system cannot find a sentence that supports a score, that score should not exist.

Regulated sectors need one more check: whether what was said about the product is correct according to the official documentation. A brilliant conversation with a false claim in it is a risk, not a success.

The trend, not the snapshot

A single session says little: the scenario, the day and the simulated customer all play a part. What matters is the trend in the same scenario across several sessions, which criteria improve and which ones stall. Only raise the difficulty of the simulated customer once the basic criteria hold up.

Biases to watch for

  • Rewarding length: talking more is not selling better. A listening criterion balances it out.
  • The halo effect: a strong opening should not inflate everything else. Each criterion is scored with its own quote.
  • Transcription errors: an accent or a noisy line can penalise someone who does not deserve it. Low scores are checked against the audio.
  • Rewarding the script: reciting the pitch is not listening. What gets scored is whether the answer fits what the customer actually said.
  • Judging the person instead of the conversation: you measure what was said, not how the speaker was feeling.

Finally, calibrate. Every so often a trainer scores a sample of sessions by hand and compares the result with the system’s. If they disagree, you adjust the anchors; you do not force the outcome.

Back to the contents

02Regulation

The EU AI Act and AI-based training: what applies now and what arrives in 2027

· 2 min read

Since 2 August 2026, people must be told when they are talking to an AI. On 2 December 2027 the high-risk obligations arrive for systems that assess workers. This is what makes sense to do now.

Applies now: say that it is an AI

Article 50 of Regulation (EU) 2024/1689 has applied since 2 August 2026. It requires that anyone talking to an AI system knows it, unless that is obvious from the context. In a training role-play the context usually makes it clear, but it is better not to rely on that: a notice at the start of every session, in writing and spoken aloud, removes any doubt. In addition, providers of systems that generate synthetic audio, images, video or text must mark them in a machine-readable way.

Already banned: inferring how the worker feels

Since February 2025, Article 5 has banned, in the workplace and in educational institutions, systems that infer a person’s mood or feelings from their voice, face or other biometric data, except for medical or safety reasons. For sales training this means something very specific: the system may assess what was said and how the conversation was handled, but it may not conclude that the salesperson was nervous, insecure or bored.

Arriving in December 2027: high risk

The Regulation treats systems intended to monitor and evaluate the performance and behaviour of workers as high risk. Following the 2026 amendment (Regulation (EU) 2026/1744), those obligations apply from 2 December 2027: risk management, technical documentation, record-keeping and human oversight, plus, for the company using the system, informing the affected workers and their representatives beforehand.

The classification depends on the intended use and is worth reviewing with an adviser. A private rehearsal whose result only the person practising can see is far from that scenario. If the scores feed into appraisals, incentives or decisions about someone’s career, the system falls squarely within it.

What to do now

  • Announce at the start of every session that the other party is an AI, even if it seems obvious.
  • Review the assessment criteria and remove any that depend on guessing how the person feels.
  • Separate private rehearsal from assessment, and decide in writing which results the company sees and for what purpose.
  • Ask each person for consent before sharing their results.
  • Start documenting the criteria, the sources and who reviews the scores now: it is the foundation of what will be required in 2027.
  • Explain to the people who use and supervise the system what the tool does and what it does not do.

This text is for guidance only and does not replace legal advice on each specific case.

Back to the contents

03Public sector

AI in public procurement: prepare, check and draft, yes; decide, no

· 2 min read

AI can save hours on every procurement file. What it cannot do is take the place of the person who signs. Where the line is, and how to design a system that respects it.

A procurement file is, to a large extent, paperwork: gathering scattered data, calculating amounts, checking that nothing is missing and drafting documents that follow well-known templates. That is where AI adds a great deal. But the file ends in administrative acts that someone has to justify and sign, and that responsibility is not delegated to a machine.

Yes: prepare

Extract from the needs report, the budget or previous contracts the data the file requires: subject matter, duration, lots and unit prices. Propose the calculation of the estimated contract value with its breakdown, so that the civil servant reviews it instead of building it from scratch.

Yes: check

Verify that every document required for the type of contract is there, for example the report on the lack of in-house resources in a services contract; that the figures match across documents and that the deadlines are consistent. If something is missing, the file must not move forward: a clear block prevents mistakes that later cost an appeal.

Yes: draft

Generate the drafts of the justification report, the internal reports and the tender specifications from the organisation’s own templates, with its name and coat of arms, ready to review in Word and Excel. The civil servant corrects and signs; the draft saves them the mechanical part.

No: decide, award or audit the spending

The need for the contract, the procedure and the criteria are decided by the contracting authority. Assessing the parts of the bids that depend on value judgements is the job of the evaluation board or the expert committee, and the award belongs to the contracting authority. Financial control belongs to the comptroller. AI can flag that a bid looks abnormally low, but ruling on it requires asking the bidder for an explanation and deciding with stated reasons. A person does that.

If an administrative action is ever automated, Spanish Law 40/2015 requires deciding beforehand which body is competent and who answers for it. When AI only prepares, checks and drafts, the person who signs remains responsible, and the system must make that clear.

How to recognise a good design

  • Every extracted data point shows which document it comes from.
  • Every draft can be edited before it is signed.
  • Blocks explain what is missing and why.
  • At the end, the complete file is exported as a package that anyone can review.
Back to the contents

04Models

Model radar, September 2026: what to watch and how to choose

· 2 min read

September brought a new generation from almost every major provider. Which families are worth following, and a method for choosing a model by task, cost and latency that will survive the next release.

The families worth watching

OpenAI has launched the GPT-6 generation, with a flagship model for reasoning and two lighter tiers for everyday work and high volume. Anthropic has refreshed its top range with Claude Opus 5.5, strong on long tasks and on agents that use tools. Google refreshes Gemini Flash every few weeks, now at version 3.8, and publishes dedicated models for live voice and text-to-speech. xAI has released Grok 4.7. DeepSeek keeps up the pressure on cost with open-weight models, and the best open models (Qwen, GLM, Kimi and others) now compete with closed models from only a few months ago. Mistral remains the leading European option.

Choose by task, not by leaderboard

General leaderboards change every week and measure things you may not need. It is more useful to sort the tasks of the business into three groups:

  • Hard, infrequent reasoning, such as analysing a contract, reviewing a long report or planning: the most capable model pays off, even if it is slower and more expensive.
  • Everyday, medium-volume work, such as drafting, summarising, classifying or scoring a conversation against a rubric: a mid-tier model usually delivers the same quality for much less.
  • High volume or instant response, such as extracting fields or talking on the phone: light or low-latency models. In voice, what matters is answering in under a second and not talking over the person.

Cost and latency are measured, not assumed

The price per token is misleading: a cheap model that needs three attempts or writes very long answers ends up costing more. What counts is the cost per task completed and done well. That is why it is worth measuring the cost of every message in production and setting a monthly cap per agent. For latency, measure the time to the first word and not just the total, especially in voice.

A method that lasts

Keep a test set of twenty or thirty real cases with the expected answer. When a new model comes out, run it, compare quality, cost and time with the current one, and decide. If switching models forces you to rebuild the agent, the problem is not the model, it is the architecture: the agent should switch models with a single setting, and it is wise to keep a second model ready in case the first one fails.

Back to the contents

05Customer service

AI voice in hospitality: what works at night and what does not

· 2 min read

Between midnight and breakfast, the front desk of many hotels depends on a single person or on nobody at all. A voice receptionist helps a lot with some things and not at all with others. This is what we learned preparing a demonstration front desk.

What works

  • Frequently asked questions at any hour: breakfast times, check-out time, parking, wifi or how to get there. They make up most night-time calls and have fixed answers.
  • Serving guests in their own language. Someone calling from abroad in the small hours appreciates being understood the first time, and the voice front desk switches language without switching numbers.
  • Taking down requests: a late check-out, a cot, an early taxi. The AI does not confirm what it cannot confirm; it writes the request down, with room, time and language, so the morning team can deal with it first thing.
  • Freeing up the night receptionist to look after whoever is standing at the desk.

What does not work and should not be attempted

  • Emergencies. A health problem, a fire or a security incident is passed immediately to a person or to the emergency number, 112. The AI does not assess or reassure in those cases: it transfers.
  • Payments over the phone. Card details are never requested from, or dictated to, an AI voice. If something has to be paid, a secure link is sent or it is done at the desk.
  • Confirming availability or prices without a connection to the booking engine. If the system cannot see the real inventory, it says it has taken note and that the request will be confirmed; it does not make things up.
  • Sensitive complaints. An angry guest in the middle of the night needs someone who can decide on compensation. The AI records the case and escalates it.

Three design rules

First: announce from the very first second that this is an AI receptionist. The EU AI Act has required it since August 2026 and, on top of that, it avoids misunderstandings. Second: always leave a way through to a person or, if nobody is available, to an on-call phone number. Third: measure. Every call should end in a structured record, with who called, what they asked for, in which language and what state it is in, so the morning team does not have to listen to recordings.

Done well, AI voice does not replace the front desk: it covers the hours when nobody answers and hands over the work in order to whoever arrives in the morning.

Back to the contents

06Privacy

Privacy by design when assessing teams

· 2 min read

If a team is afraid of what the company will see of its rehearsals, it will practise less and worse. Privacy does not slow training down: it is the condition for it to work.

Consent, person by person

Each person decides whether to share their individual results with the company, and can change their mind. The question is asked clearly, separately from other notices and with no consequences for saying no. Until there is consent, that person’s results do not appear under their name in any report, and if they withdraw it, the results stop being shown from that moment on.

This is the direct application of the GDPR principle of data protection by design (Article 25): privacy is the starting point, and sharing is an active decision.

Small groups are not shown

An aggregated report can give a person away if the group is small. If a team has three members and two of them share their scores, the average reveals the third. That is why groups of three people or fewer are not shown in the consolidated reports, and anyone who has not given consent only counts towards totals that cannot identify them.

Personal rehearsals are private

Some rehearsals only serve the person doing them: preparing for a difficult conversation, a negotiation or an interview. Those rehearsals never reach any company dashboard. If the person wants to show them, that is their decision. Someone who knows their rehearsal is private dares to practise what they find hardest, which is exactly what they need most.

What is measured and what is not

It is the conversation that is assessed, not the person: what was said, whether the methodology was followed and whether the facts were correct. Nobody’s feelings are inferred from their voice, something the EU AI Act bans in the workplace. And the criteria are visible to whoever is being assessed: nobody should receive a score based on rules they do not know.

Checklist

  • A written purpose: what the results are used for and what they are not used for.
  • Individual consent that is revocable and recorded.
  • A minimum group size in every aggregated report.
  • Role-based access: the trainer sees what they need in order to train, and management sees trends.
  • A retention period for recordings and transcripts that is defined and enforced.
  • The right of every person to see their own sessions.
Back to the contents

If one of these pieces describes your case, let’s talk.

We will show you how we solve it in Fenix, with a case similar to yours.