← All chaptersChapter 8 of 8

From workflow to responsible implementation

You bring a workflow into controlled use, measure real value, and know when not to automate.

After this chapterYou can design and run a bounded practice pilot with fictional data, human checks, measurement points and a supported decision to test further or stop. You identify the additional evidence still missing for real deployment.
Your progress0 of 40 lessons
8.1

Standardize before automating

Automation accelerates a stable process, but also increases existing uncertainty and errors.

First map out the current way of working: trigger, input, decisions, exceptions, output, and owner. Standardize what must go well and define where human expertise remains necessary. Then automate only predictable, low-risk steps.

Use three levels: assistance produces a draft, semi-automation performs fixed operations with checks, and automatic action is appropriate only when impact is low and error handling is strong. Expert covers the design and evaluation of integrations; a complete technical implementation requires additional development knowledge.

  • Process first
  • Exceptions visible
  • Assistance versus action
  • Low impact for autonomy
  • Owner still needed

Terms in plain language

API
A technical interface through which software passes information or instructions to other software.
How you can use this

A support workflow may classify requests and create a draft; reimbursement and contract changes require separate authority.

Try this prompt
Map [process] and classify each step as human, an AI draft, semi-automatic or a candidate for automatic action. Explain your choices in terms of impact, uncertainty, recoverability and exceptions.
Knowledge check

You want to partly automate a recurring study schedule. Three known exceptions have no agreed handling yet. Which design step comes first?

Your practical exercise

Create a process map and delete any automation for which there is no clear error handling.

8.2

Human approval as a design component

Define who can approve what, with what information and within which time frame.

A simple approval button is not control. Show the proposed action, used source data, changes, risk, rollback option, and deadline. The approver must be authorized and sufficiently informed.

Determine behavior in case of no response: remind, escalate, safely stop, or follow a low-risk standard path. Never carry out a risky action because someone did not respond on time. Log proposal, decision, and execution result proportionally to the risk.

  • Authorized approver
  • Decision information
  • Deadline and escalation
  • Safe default
  • Recoverability

Terms in plain language

Logging
Recording which relevant action took place and what its outcome was, without unnecessary sensitive content.
Escalation
Passing a problem to someone with the right knowledge or decision-making authority.
How you can use this

An external quote shows changed amounts, conditions, and source data before sending can be approved.

Try this prompt
Design an approval gate for [action]. Specify authority, information needed for the decision, risk, deadline, reminder, escalation, safe default, logging and rollback method.
Knowledge check

A reviewer must approve a message containing changed data before it is sent. There is no authorised replacement. The deadline approaches and the reviewer does not respond. Which handling fits this approval gate?

Your practical exercise

Design two approval gates: one for low and one for high impact.

8.3

Privacy, rights, and transparency in the workflow

Classify and minimize data before input; do not treat pseudonymization as anonymization.

Use entirely fictional data for these exercises. Publicly available personal data are not automatically free to use. Internal information requires consideration of policy and necessity; passwords, API keys, identity documents and complete sensitive case files do not belong in an ordinary prompt. Replacing names with codes is usually pseudonymisation: links may remain through other characteristics.

Check rights to source material and output. Document your own creative choices. Some uses are subject to transparency obligations under Article 50 of the European AI Act, for example direct interaction with AI or deepfakes. Which obligation applies depends on your role, the application and any exceptions. Check the current official explanation before deploying such an application.

  • Classify data
  • Minimize
  • Correctly naming pseudonymisation
  • Check rights
  • Transparency where needed
How you can use this

A customer case used for learning is made entirely fictional; an internal dossier is not declared safe merely by replacing a name with “Customer X”.

Try this prompt
For [an entirely fictional description of a workflow, with no real files or personal data], create a data and rights register: category, necessity, safe placeholder, retention needs, access, source rights, transparency and stopping rule. Do not describe pseudonymised data as anonymous.
Knowledge check

You are testing a document review. The real dossier contains personal data; you can make the same relevant structure and exceptions entirely fictional. Which input best fits this course exercise?

Your practical exercise

Rework a fictitious risky scenario into minimal, safe training input.

Source for this lesson

European Commission – AI transparency
Article 50 has conditions of application and exceptions; there is no general obligation to label every text as AI-generated.
Checked: 2026-09-07

8.4

Measuring value and quality

Measure time, correction work, usability, and risk; prompt volume is not a business value.

Conduct a baseline measurement before the pilot. Record lead time, error or correction rate, quality level, and usage frequency of the existing process. Measure during the pilot using the same definitions. Add a knock-out measure for critical errors or privacy incidents.

A draft produced faster but requiring extensive rework is not a gain. Measure the total effort needed to reach an approved result. Use a small, representative sample and report uncertainty. A positive pilot does not yet prove scalability to other teams or data.

  • Baseline measurement
  • Same definitions
  • Total corrective effort
  • Critical error separate
  • Limited conclusion

Terms in plain language

Baseline measurement
Measuring performance before the change using the same definitions you will use later.
How you can use this

A report workflow measures minutes to approval, number of factual corrections, rubric score, and incidents; not just generation time.

Try this prompt
Design a measurement plan for [pilot] with baseline measurement, numerator, denominator, data source, sample, quality rubric, critical error, measurement period, and scale decision. Avoid unproven productivity claims.
Knowledge check

Without AI, a completed task takes 40 minutes. With AI, generation, correction and final review take 7, 25 and 10 minutes respectively. Which comparison uses the same endpoint?

Your practical exercise

Use three different fictional study-club messages. First assign each a checked label manually; measure the time including your check. Then perform the same task with your prompt and measure again through to a reviewed result. Done: record both times, correctness and corrections for each case. State that recognising the same cases may favour the second round; three practice tests do not prove business gains.

8.5

The advanced final project

Demonstrate mastery with a working, tested workflow and a fair implementation decision.

Choose a bounded task and create a fictional dataset for this exercise. Use real work data only in an authorised environment and in accordance with applicable policy. Provide a result definition, context register, phases, prompts, source checks, test set, sample output, human approval and measurement plan. Keep genuinely failed tests and the corrections they produced. If you find no error, add an incorrect answer clearly labelled as deliberately altered to test your checking process; do not present it as model output.

Finish the exercise with a decision: continue limited testing, redesign or stop. Real deployment needs a separate authorised decision and appropriate business evidence. Support this with quality, time, risk, maintenance and transferability. Identify which technical integration or business governance is addressed only in Expert.

  • Working workflow
  • Source and privacy control
  • Test evidence
  • Measurement plan
  • Scale, test, or stop decision
How you can use this

Possible projects: research file, document analysis, editorial workflow, team update, or controlled plugin process.

Try this prompt
Help me plan my final project for [task]. Create an evidence checklist for goal, sources, phases, prompts, controls, privacy, tests, human approval, baseline measurement, and implementation decision. Do not execute anything yet.
Knowledge check

Your final project handles three ordinary examples well. Two predefined edge cases fail; their impact and whether they can be fixed have not been assessed. Which final decision fits this evidence?

Your practical exercise

Carry out the project with fictional data and support all seven rubric criteria with your dossier. Another user may assess it. Working alone: close your first assessment, reopen the source and rubric and reassess each criterion in a separate round. Sufficient means that the requested evidence is present and the checks pass; open risks and business steps not performed remain explicit. Label this self-assessment and limit your final decision to further practice, redesign or stopping this practice version.

Worked example

Compare a baseline and pilot fairly

Fictional practice material; incorrect answers have been created deliberately for this exercise.

An entirely fictional measurement case for study sheets. Three existing tasks take 36, 40 and 44 minutes to produce an approved sheet. Quality criteria are set in advance: definitions match the source, exercises can be answered, and answers are separate. An invented source claim is a critical error.

Input

Measure three comparable, new source packages using the same criteria. For each sheet, record generation time, corrections, the final check and critical errors. Record 24 minutes of one-off setup separately and also include them in the total for this first pilot. Record actual work; do not fill in favourable estimates for missing times.

First practice answer

Fictional pilot record in minutes: sheet 1: 6 + 18 + 4 = 28; sheet 2: 7 + 20 + 4 = 31; sheet 3: 8 + 22 + 7 = 37. The flawed report counts only the average 7 minutes of generation and claims a saving of 33 minutes.

Check

Baseline: (36 + 40 + 44) / 3 = 40 minutes. Pilot execution: (28 + 31 + 37) / 3 = 32 minutes. Including setup: (96 + 24) / 3 = 40 minutes. In the fictional record, 3 out of 3 final sheets meet the criteria; 0 critical errors were observed. The latter does not prove zero risk.

Improved version

For these three tasks, no overall time saving has been demonstrated when setup is included. Execution alone is 8 minutes shorter on average. Decision: continue limited testing with new source packages; set in advance how many comparable tasks will follow and use the same final criteria. Report differences in source length and difficulty.

Try it yourself

The next sheet takes 9 minutes to generate, 25 to correct and 8 to check. What execution time do you record, and what do you conclude compared with 40 minutes?

View the model answer

42 minutes, so this task takes 2 minutes longer. Add the observation; do not hide it behind the earlier average. Assess quality, maintenance and greater variation before recommending the method more widely.

Chapter assignment

Bring everything together

Complete a fictional practice pilot with a baseline, workflow, test set, example result, correction log, privacy and rights register, and approval gate. Conclude with an evidence-based decision: continue limited testing, redesign or stop. State which additional decision by an authorized person and business evidence are still needed for real implementation.

Maximum 10,000 characters per note.

Progress and notes are stored only in this browser on this device. Do not enter sensitive data. Download your notes regularly. This course sets no automatic expiry date. You can delete the data through your browser’s site-data settings; export anything you wish to keep first. Browser settings or cleanup may erase it earlier. These local notes are not sent to Finaudax.

My assessment

Marking a lesson complete is your own assessment; it does not automatically demonstrate mastery.

Result

Not yet sufficient: The intended use or testable requirements are missing.

Sufficient: The final product demonstrably meets a realistic use case.

Strong: Another user can assess the result against the same requirements.

Process

Not yet sufficient: There is only a long prompt; intermediate results and stopping points are missing.

Sufficient: Phases, intermediate products, stop conditions, and responsibilities are explicit.

Strong: Someone else can carry out the process and knows what to do when an error occurs.

Sources

Not yet sufficient: Sources and assumptions cannot be distinguished.

Sufficient: Source facts, assumptions and interpretations remain traceable and separate.

Strong: Conflicting and missing evidence is also recorded in a traceable way.

Test evidence

Not yet sufficient: There is only one successful example, without criteria set in advance.

Sufficient: Normal, difficult, and unsafe cases were tested with predetermined criteria.

Strong: New cases and repeated runs have been assessed in addition to the development tests; critical errors are reported separately.

Safety

Not yet sufficient: Data or external actions have not been checked for permission and risk.

Sufficient: Privacy, rights, prompt injection, and human approval have been appropriately addressed.

Strong: Unsafe fictional input and the stopping or recovery path have also demonstrably been tested.

Improvement

Not yet sufficient: A change is called better without comparable evidence.

Sufficient: At least one change is better substantiated with comparable evidence.

Strong: The improvement holds up on new cases; limitations and maintenance are clear.

Pilot and decision

Not yet sufficient: A baseline, total handling time including review and rework, or a reasoned decision to implement, test further or stop is missing.

Sufficient: A comparable baseline and pilot measure quality and total handling time. The decision follows predefined thresholds; limitations and next steps are stated.

Strong: New cases confirm the findings; variation and exceptions are visible. Another person can reconstruct the decision and knows when reassessment or stopping is necessary.