← All chaptersChapter 2 of 8

Risk, data, and solution choice

You choose the working method and model only after data, error impact, rights, and operational requirements are clear.

After this chapterYou can classify a use case, choose a suitable architecture, and document risks and permissions.
Your progress0 of 48 lessons
2.1

Privacy, copyright, confidentiality, and hallucinations

Treat four risk domains separately; one general warning is insufficient.

Privacy is about personal data; confidentiality about business information; copyright about protected material; hallucinations about invented or incorrectly linked content. One use case can affect all four.

For each risk, document the scenario, likelihood, impact, prevention, detection, owner and recovery. Pseudonymisation sometimes reduces risk but does not automatically make data anonymous.

Before a real pilot, add a legal starting check: which application, which organisation and role, which people, which data and which consequences? Check the AI Act for prohibited uses, possible high-risk classification and transparency. A product name or low internal risk score does not answer these questions. The provider places the system on the market or puts it into service under its own name; a deployer uses it professionally under its authority. Have any unclear classification assessed before the relevant practical trial.

In Noor’s ordinary text pilot, the customer stays outside the AI chat: a staff member reviews and sends the reply personally. A bot that talks directly to customers is a different application. Article 50 may then require information about the AI interaction. For publication, relevant cases include deepfakes and certain AI texts on matters of public interest; substantive human review and editorial responsibility may provide an exception for those texts. One general label does not automatically cover every situation or provider obligation.

As of 8 September 2026: the AI Omnibus has changed the timeline. Article 50 generally applies from 2 August 2026; certain existing systems have a transition until 2 December 2026 for the provider obligation under Article 50(2). The main rules for high-risk systems in Annex III follow on 2 December 2027, and those for high-risk AI embedded in regulated products in Annex I on 2 August 2028. Use the official timeline for a real starting decision; existing privacy and employment rules remain relevant in the meantime.

  • Personal data
  • Trade secret
  • Usage rights
  • Factual error
  • Recovery
How you can use this

A quote assistant processes contact data, secret pricing rules, protected source text, and possibly invented terms.

Try this prompt
Use a completely fictional description of a process, without real case files, personal data or secrets. Build four risk registers for [use case]: privacy, confidentiality, copyright and incorrect output. Include prevention, detection, ownership and recovery.
Knowledge check

A quote assistant uses contact details, internal pricing rules and supplied product descriptions. It may also add an incorrect guarantee. Which risk analysis is appropriate?

Your practical exercise

Analyze one data flow and assign an owner for each risk.

Source for this lesson

European Commission – Transparency under Article 50
Provider and user roles, direct AI interaction, deepfakes and public texts; conditions and exceptions differ.
Checked: 2026-09-08

European Commission – Current AI Act timeline
Application dates following the AI Omnibus, including the limited transitional regime for Article 50(2).
Checked: 2026-09-08

Official Journal – AI Omnibus 2026/1744
Amending regulation; read alongside the AI Act and current official implementation information.
Checked: 2026-09-08

2.2

Classify data before use

Access to data does not automatically mean AI processing is permitted.

Assess two things separately: how confidential is the information, and does it contain personal data? A public document can also contain personal data. For each field, document why you need it, which contractual or sector-specific rules are relevant, and which environment, access and retention period are permitted.

Use entirely fictional information in these exercises. Synthetic data is artificially generated but may still create risks if it retains recognisable details from real case files. Simply labelling information synthetic or replacing a name does not make it anonymous. Classify the required data types before sending the content to a service.

The GDPR, called AVG in Dutch, applies to personal data. Record the purpose, legal basis, roles, minimum data and information for the people concerned. Before use, assess whether a data protection impact assessment is needed: a DPIA, or GEB in Dutch. It is mandatory when the processing is likely to pose a high risk to rights and freedoms. Consult the Belgian Data Protection Authority’s criteria and mandatory list; new technology alone does not make every exercise require a DPIA. Document the conclusion and involve the DPO or an appropriate specialist. If a high residual risk remains after the measures, prior consultation with the competent supervisory authority is required.

  • Classification
  • Purpose limitation
  • Minimum fields
  • Retention
  • Access

Terms in plain language

Retention
How long data is actually kept and under which rules.
AVG / GDPR
The European General Data Protection Regulation; GDPR is the English abbreviation and AVG the Dutch one.
GEB / DPIA
A prior assessment of privacy risks and measures for a proposed data processing activity.
DPO
Data protection officer; an independent advisory and monitoring role where one is appointed.
How you can use this

A service report uses problem category and solution, but omits name and full file when they are not needed.

Try this prompt
Use a completely fictional description of a process, without real case files, personal data or secrets. Create a data classification for [workflow]. For each field, provide its purpose, category, necessity, environment, retention and a fictional test alternative.
Knowledge check

An employee wants to use a public list of names and personal email addresses to test a new AI workflow. Fictional records would suffice. What is appropriate?

Your practical exercise

Remove every field without a demonstrated need. Also record in row H2 of your pilot worksheet: personal data yes/no with a reason, the intended purpose and legal basis, who assesses the roles and the need for a DPIA, and what evidence is still missing. Work on paper with fictional data; an unresolved starting condition becomes an action with an owner, not tacit permission.

Source for this lesson

EDPB – AI models and GDPR
Publicly available personal data is not automatically exempt from privacy requirements.
Checked: 2026-09-07

Belgian Data Protection Authority – Data protection impact assessment
Assessing and documenting in advance whether a DPIA is needed; mandatory cases and high residual risk.
Checked: 2026-09-08

EDPB – Controller or processor
Roles and the associated GDPR responsibilities; more than a product setting alone.
Checked: 2026-09-08

2.3

Model choice per task

Choose based on task requirements and evaluation evidence, not on reputation or a single demo.

Establish modalities, tools, context, accuracy, and output format. Test writing, analysis, code, research, and images with different representative datasets where their requirements differ.

Model names change. Keep requirements, the test set and the minimum score as a stable foundation, and retest when the model or prompt changes.

  • Task fit
  • Modality
  • Tools
  • Quality threshold
  • Retest

Terms in plain language

Modality
The form of information, such as text, images or sound.
How you can use this

An imaging workflow and financial extraction process use different tests and possibly different models.

Try this prompt
Create a model selection protocol for [tasks]. Define requirements and test set; compare quality first and then speed and cost. Do not name a winner without measurement results.
Knowledge check

Two models pass the normal tests for a quotation task. In a critical test, the cheaper model invents a missing contract term. The other option marks it as unknown. How do you choose?

Your practical exercise

Create five normal and two critical fictional tests with expected behaviour defined in advance. Test two available options if you have access to them. If your free account only lets you use one option, test it and leave the second column marked ‘not performed’. Completion check: the source, criteria and your own outputs are saved; you do not choose a winning model without comparable observations. Purchasing access to more models is not a course requirement.

Source for this lesson

OpenAI – Model selection
Task requirements and evidence of quality come before optimisation; this source covers the API.
Checked: 2026-09-07

2.4

Speed, accuracy, context, and cost

Optimize only after minimum quality and safety limits have been met.

Latency affects usage; context determines available information; requests and tokens influence time and costs. More context can also add irrelevance and conflict.

Measure end-to-end time, corrective work and the cost of errors. Shorter output, fewer requests or a lighter model are valid only when the evaluation set confirms that quality is maintained.

With a p95 measurement, approximately 95% of the measured completion times are at or below the stated value. Also document the calculation method: software packages may interpolate percentiles differently. For a small practice set, you can sort the times and round the position up from 0.95 × the number of measurements. For 20 times, this gives the 19th time. Twenty measurements do not establish a stable performance guarantee.

  • Minimum quality
  • End-to-end time
  • Context
  • Volume
  • Error costs

Terms in plain language

Latency
The wait for a system response; not necessarily the total completion time.
Request
A single request sent to a system.
Token
A unit in which a model processes text or other input; it is not the same as one word.
p95
A percentile describing the slower end of measured completion times.
Evaluation set
The examples and criteria used to test quality.
How you can use this

An answer that is generated one second faster but needs correction more often may take more time overall. Therefore measure the complete process, including review and corrective work.

Try this prompt
Design a matrix for [workflow] with minimum quality, p95 lead time, context requirement, volume, error costs, and budget. Propose three optimizations to test.
Knowledge check

A shorter answer format reduces generation time, but employees more often have to look up missing conditions. What should an optimisation test compare?

Your practical exercise

Extension: time twenty fictional tasks you carry out, including correction and rechecking. Basic route for the calculation: use twenty simulated durations in minutes: 4, 5, 5, 6, 6, 6, 7, 7, 7, 7, 7, 7, 8, 8, 8, 9, 9, 10, 12, 15. Completion check: the stated method gives p95 = 12 minutes; say whether the durations were measured or supplied. These durations do not belong to Noor’s pilot and do not prove model performance.

Source for this lesson

OpenAI – Model selection
Task requirements and evidence of quality come before optimisation; this source covers the API.
Checked: 2026-09-07

2.5

ChatGPT, Work, plugin or API

The interface choice determines data flow, scale, review, and management.

Chat suits a focused conversation; Work suits larger deliverables that can be reviewed; plugins suit connected data or actions; and an API suits integration into your own product, scale or systematic logging.

Availability and permissions differ. Document identity, source rights, approval moment, and storage for each route.

  • Experience
  • Sources
  • Actions
  • Scale
  • Logging
  • Permissions
How you can use this

A periodic internal analysis can fit in Work; real-time processing in a client portal requires more of a managed integration.

Try this prompt
Compare Chat, Work, plugin, and API for [use-case] on data, identity, rights, review, logging, scale, and management. Mark unconfirmed availability.
Knowledge check

An SME wants to include AI in its own customer portal, with user identity, controlled permissions and an audit trail for each request. Which solution route do you investigate specifically?

Your practical exercise

Draw the data and authorisation flow for the chosen option.

2.6

Recording an architectural decision

Also record why you choose, under which assumptions, and when you reconsider.

Describe context, requirements, examined alternatives, choice, consequences, and owners. This way, a later team does not treat the choice as a law of nature.

Add review triggers such as growth in volume, new sensitive data, a model change, an incident or a supplier change. Link the decision to tests and the risk register.

  • Context
  • Alternatives
  • Decision
  • Consequences
  • Proof
  • Trigger

Terms in plain language

Architecture decision
A short document describing the chosen technical approach, alternatives and reasons.
How you can use this

An API choice is reviewed when volume doubles or a new data category appears.

Try this prompt
Write an architecture decision for [use-case] with context, requirements, options, choice, consequences, evidence, owner, and review triggers.
Knowledge check

An architecture decision allowed only internal product information. The new workflow version must also process customer files. What do you do with the existing decision?

Your practical exercise

Assess the architecture decision from the perspectives of technology and business use. Real owners can review it; solo, take on these two fictional roles separately. For each role, record a requirement, evidence and an unresolved condition. Completion check: unconfirmed access or data processing is not approved just because a role has been assigned on paper. Complete row H2 of the pilot worksheet with the data flow, four risks, chosen route and unresolved conditions; label the review as self-assessment.

Worked example

What data does the draft answer need?

Fictional practice material; incorrect answers have been created deliberately for this exercise.

This case uses only invented people, fields and systems. Atelier Noor is preparing draft answers. The chosen service and its retention policy still need assessment; this exercise makes no statement about a real account.

Input

A fictional customer enquiry contains: name Lina Voorbeeld, email address lina@example.invalid, order number DEMO-104, product code LAMP-2, the question “Is this lamp dimmable?”, an invented medical explanation and a link to a public profile of Lina. The approved practice sheet says: “LAMP-2 is not dimmable.” The AI only needs to draft a general answer; it does not look up an order or send anything.

Deliberately flawed practice answer

“Replace Lina with Customer X and upload the entire enquiry. The profile is public and therefore freely usable. The context disappears automatically afterwards.”

Check

A different name does not automatically anonymise the remaining data. The product question does not require an email address, order number, medical explanation or profile. Public information about a person remains a separate privacy question. The end of a task does not prove that chats, files, technical logs or copies held by other services have been deleted.

Improved result

The practice input consists of product code LAMP-2, the general question and the approved sheet. The employee receives the draft and links it back to the right customer outside the AI. The data flow is: selected practice fields → assessed AI service → draft → human review. Before real use, the responsible person must assess the permitted purpose, provider, settings, access and deletion route. An internal proposal to “keep trial logs for no more than 30 days” is a design choice that still needs to be assessed and implemented; it is not a promise about the service. Where possible, log only necessary technical metadata. Check whether that metadata still refers to a person.

Try it yourself

The next fictional question concerns the general warranty period for LAMP-2. Available fields are product code, purchase date, home address and bank account number. No assessment of an individual warranty claim is needed. Which fields do you send to AI, and what else do you consult?

View the model answer

For this general explanation, use the product code and approved warranty terms. Purchase date, home address and bank account number are not needed here. An individual claim would be a different task with different necessary data and checks.

Chapter assignment

Bring everything together

Deliver data classification, four-part risk register, model test plan, channel choice, and architecture decision.

Maximum 10,000 characters per note.

Progress and notes are stored only in this browser on this device. Do not enter sensitive data. Download your notes regularly. This course sets no automatic expiry date. You can delete the data through your browser’s site-data settings; export anything you wish to keep first. Browser settings or cleanup may erase it earlier. These local notes are not sent to Finaudax.