BLOG – Blog de Fermin Fernandez / Sat, 08 Aug 2026 21:26:22 +0000 en-US hourly 1 https://wordpress.org/?v=7.0.3 /wp-content/uploads/2023/05/cropped-IconoFF-32x32.jpg BLOG – Blog de Fermin Fernandez / 32 32 Generative AI: Reflections at the end of 2025 /en/generative-ai-reflections-at-the-end-of-2025/ Tue, 30 Dec 2025 21:41:00 +0000 /?p=311 Continue readingGenerative AI: Reflections at the end of 2025]]> GenAI Reflections

2025 is coming to an end, and it’s a good time to reflect on the highlights I’ve experienced as a consultant in the solutions automation sector. While electronic invoicing solutions in Europe deserve a dedicated chapter (I promise to talk about that in 2026), today I want to focus on Generative Artificial Intelligence. Without a doubt, it has been this year’s protagonist, taking center stage in most of my meetings, training sessions, debates, and challenges.

Generative AI: Between the Hype and Reality

Every week we can read a new post, study, or promise about Generative AI. Isn’t it getting a bit overwhelming? There’s also a palpable sense of urgency from many companies to jump on this bandwagon, sometimes without a clear destination, and with the (almost threatening) message that “if you don’t invest in AI, you’ll be left behind.”

It’s clear that Generative AI has made our daily work easier: improved reports, more polished presentations, clearer emails, quick access to information that used to be hard to find… especially useful when you know how to make the most of it. The use of Gen AI services keeps growing, and it’s undeniable that they’re valuable tools for many professionals.

However, when our clients consider investing in Generative AI projects, the perspective changes: what truly matters is the return on investment (ROI), whether through cost reduction, productivity increases, or regulatory compliance. It’s not enough to have AI for AI’s sake; it must provide clear and tangible value. And curiously, there aren’t that many publications about  these concrete, real-world cases.

AI Agents and Real ROI

This is where the much-discussed “AI Agents” come in: solutions that automate specific processes using AI engines, able to make decisions and free users from repetitive tasks. That’s where the real ROI lies.

But it’s not easy, and there are many challenges. Many of you have probably heard about the MIT study published in August, which indicates that most Generative AI pilots fail. Are you aware that most people commenting on this report haven’t actually read it?

To put it in context, it’s not a technical study or in-depth research report, but rather an analysis based on interviews with a limited number of companies, and it considers a “failure” to be not achieving clear ROI within a few months. Personally, I find that a pretty strict criterion for considering a project as failed. It’s important to note that this doesn’t necessarily mean the projects didn’t work, but that they didn’t generate a clear benefit for the company within that timeframe.

Still, the underlying message is valid: many AI pilots don’t make sense because they’re often launched without clear objectives, simply to justify investment in an AI department. Other times, they overlooked that business processes are complex and full of interactions, so focusing only on Generative AI isn’t enough—you need to combine Gen AI with other types of automation applications as well.

Generative AI Is a Tool, Not the Center of the Universe

In meetings with large companies (IBEX35), I always wonder: does it make sense for them to have independent AI departments? What’s clear to me is that Generative AI cannot automate processes on its own. It needs to be integrated with the rest of the company’s ecosystem (web services, databases, RPA, ERPs, CRMs, email, etc.), manage exceptions with users, and coordinate with other business processes.

That’s why I see Generative AI as just one more tool within a broader set of automation solutions. It’s not a solution by itself, but a component that, when used well, enhances everything else.

Many companies already use process automation solutions that have proven effective. You simply have to add the Generative AI component needed for each process. Most of these solutions already include Gen AI integrations. That’s why I believe the most sensible approach is to continue relying on the teams responsible for Process Optimization (whether they’re called Operational Excellence, Continuous Improvement, Digital Transformation, or something else) and train them in Generative AI, so they can make the most of all available tools.

Something similar happened in the past with the arrival of RPA technology: large companies created dedicated departments focused exclusively on that technology. Today, almost all of those have disappeared and have been integrated into process optimization teams.

Looking Ahead to 2026

The main challenge of Generative AI is its non-deterministic nature. This means that even when you ask several times the same question (same input parameters), it might not always generate the same answer. This characteristic makes automating processes or repetitive tasks difficult, as the lack of consistency can complicate integration into workflows that require predictable results.

That’s why in 2025, most projects have focused on applying Generative AI in specific departments, where consistency isn’t as important; Customer Service (handling support cases, managing information requests, gathering feedback), Marketing (content creation and market research), and IT (software development, support, and cybersecurity).

However, the core business of companies and the majority of repetitive tasks (back office)—where automation typically brings the most value—are not found here. For this reason, I believe we need to change our mindset and accept a certain degree of indeterminacy when using AI, rather than striving for total automation. I think that automating 75–80% of these processes could already provide a significant return on investment. Of course, it’s essential to properly manage the remaining 20–25% of incidents, possibly connecting the process to humans. I hope that next year, we can make the leap to these kinds of solutions!

In the meantime, I still think expectations around Generative AI remain too high. Once they become more realistic, we’ll be able to appreciate its true value: being just another tool—a very powerful one—for automating processes and improving efficiency in different areas.

Wishing you all a great 2026, full of useful projects, less hype, and more concrete results! 🚀

]]>
Why Don’t We See Generative AI in the Office? /en/why-dont-we-see-generative-ai-in-the-office/ Thu, 11 Apr 2024 09:59:00 +0000 /?p=319 Continue readingWhy Don’t We See Generative AI in the Office?]]>

If you’ve ever wondered why everyone uses generative artificial intelligence on a personal level, but there are so few use cases in the business world, here are three reasons that can explain it:

It’s too generic: Generative AI has learned from millions of data points, but this data is general—collected from thousands of public sources. However, a company makes decisions based on its own business data, which is private. These data have not been learned, so the AI doesn’t provide good answers to questions about each specific business.

It hallucinates: Generative AI produces nonsensical answers from time to time—this is known as “hallucinations.” Without the necessary business context, hallucinations are much more frequent when questions involve business-specific cases.

It’s outdated: Generative AI is not up to date. It has learned from a lot of data, but from the past, since collecting new information is a massive effort. This means it doesn’t take into account information from the last few weeks or months.

How do we solve this?

The most obvious option would be to create a custom model for each business, training the AI with our own data. The problem with this approach is that it is too costly in terms of time and money.

That’s why nowadays I’m seeing more and more solutions that use RAG (Retrieval Augmented Generation), which is a way to personalize the general generative AI model by giving it the right context before it answers our questions.

To do this, we first collect the relevant business documents, index them, and store them as vectors in a database. Then, when asking our questions, we instruct the generative AI to understand what we want, but to answer using only the information from the stored documentation. We could also point directly to the documents at the same time as we ask about them.

With this technique, we completely eliminate hallucinations because we’ll get answers based on our documents. Moreover, the solution will always be up to date, since we keep indexing new documents as they are generated.

]]>
The Qualities of a Good Pre-Sales Consultant /en/the-qualities-of-a-good-pre-sales-consultant/ Mon, 08 Feb 2021 10:08:00 +0000 /?p=321 Continue readingThe Qualities of a Good Pre-Sales Consultant]]>

Within the sales process of a software solution, you can find several roles that participate in one way or another at different stages. Of course, the main player is the salesperson, but there is also involvement from telemarketing, marketing, pre-sales, professional services, sales support, or legal department.

Pre-sales consultants are familiar with the software being offered and take part in activities such as technical presentations, demonstrations, proof of concept, responses to requests for information, and so on.

Twenty years ago, it was very common to look for pre-sales consultants who could thoroughly learn the software, so there was a tendency to seek candidates with strong technical skills; it was even better if they could program, had experience with databases, web-services, or network protocols.

But is this still the case today? In this video, I share some thoughts on the topic.

]]>
RPA & Proofs of Concept /en/rpa-proofs-of-concept/ Mon, 03 Feb 2020 10:16:00 +0000 /?p=325 Continue readingRPA & Proofs of Concept]]>

During my career, I have participated in many Proofs of Concept (POCs) using various software solutions. POCs are a resource sometimes used during the sales phase to demonstrate a series of specific, pre-agreed functionalities.

Since a POC can require significant effort without any guarantee it will result in a sale, we (software vendors) only undertake them when we see a clear commitment to purchase from a client who wants to confirm that the solution will work with their particular business requirements. As a general rule, we do not charge for this work, as our goal is to sell the solution.

A clear alternative to POCs are references, which can be used to verify that the solution works for other clients. We also often do personalized demos or even technical workshops. There are also what we call Pilot Tests, which are real installations on a small scale and in the client’s environment. These tend to last longer and incur associated costs. One could write a whole article just about these presales tools.

The fact is, the emergence of RPA technology has truly revolutionized digital transformation projects, and it’s something many companies are interested in, as it is a quick and simple way to start automating tasks. For example, in 2019 we averaged two to three new inquiries per week specifically about this technology. If we assume a similar or greater trend among our competitors, and add to this the fact that some RPA vendors now have more and more distributors (some solutions offer free online training or open-source versions), a market is being created in which almost anyone can sell you an RPA solution. Moreover, it’s becoming customary to do a POC for every interested client. That’s why today, the words “RPA” and “POC” are more closely linked than ever. And I wonder why.

Personally, I think it’s all related to the new distributors popping up everywhere. Since they don’t have good references yet, the POC is a good way to earn the client’s trust. In theory, it’s an easy technology to implement. Also, almost all of them charge for their services, so they have little to lose and, at the same time, get training on the tool by implementing a project. Naturally, this means that some POCs don’t turn out well, but it doesn’t matter much since there’s another one waiting the next week. I get the impression that more POCs are being sold than actual final products.

This way of doing business means that if a potential client likes your solution, they’ll end up requesting a POC, because others have offered them one.

Let’s be honest: any of the best-known RPA solutions are capable of automating a specific task. It may take more or less work, but it can be implemented. So, are RPA POCs really useful?

Considering that the AIIM Emerging Technologies Report (2019) indicates that only 3% of companies have managed to scale RPA technology with 50 or more robots, it’s clear that focusing solely on doing a POC for a specific task doesn’t provide the full picture.

In fact, most of the time, only that same task or perhaps one more ends up being automated. This is because the chosen solution doesn’t allow work at a corporate level, or the vendor doesn’t have the experience to help set up a center of excellence or expand the solution.

Therefore, my conclusion would be that a Proof of Concept on a specific task does not always indicate the suitability of a solution. One must also consider how far the product can go at a corporate level and who is implementing it.

]]>
Capturing Unstructured Documents /en/capturing-unstructured-documents/ Sun, 13 Oct 2019 10:09:00 +0000 /?p=332 Continue readingCapturing Unstructured Documents]]>

In the world of document capture, documents have traditionally been divided into three types: structured, semi-structured, and unstructured.

More than 20 years ago, the first capture programs focused on extracting information from structured documents, where the position of each data field was known. At that time, when web pages were not yet widespread, it was common to use this type of form to communicate with companies, and thousands of capture processes of this kind were automated (orders, surveys, work reports, exams, etc.).

Later, the automation of semi-structured document capture began, where data varies in position but rules can be created to locate them. The typical project was the capture of supplier invoices, where around 70% of the data could be captured automatically. Subsequently, the technique of learning supplier templates was introduced, now advertised as “Machine Learning.” This technique allows the software to learn the layout of each supplier’s invoice after the user indicates where the data is. Once the main suppliers are learned, and by combining learning with the use of rules, it became possible to capture more than 90% of the invoice data. Over the last 10 years, we have implemented hundreds of such projects, and today we continue to help companies that still haven’t automated their accounts payable processes.

It wasn’t until a few years ago that work began on capturing unstructured documents, where there are no rules to apply and template learning techniques cannot be used because documents of the same type (templates) are not usually repeated. This is the case for capturing data from mortgages, deeds, notarial acts, or emails received in a mailbox.

To extract information from these documents, artificial intelligence techniques must be used, and more specifically, Natural Language Processing (NLP). There are many services and programs that use NLP (Microsoft, Google, etc.), but most of them focus on extracting tags or attributes (for example, getting all the names that appear in a document) and sentiment analysis (whether the text expresses something positive or negative). Moreover, their knowledge bases are created with natural language (the language we usually speak), which is not typically the language used in business documents. That’s why it is important to use a tool capable of using NLP techniques to find specific data in a document (for example, if there are several names in a document, it should indicate who is the Notary and who is the Appearing Party). More generalist tools are usually not suitable for this purpose.

To understand how this type of solution works, here is an example of how data capture from Deeds is performed. The process is divided into three parts:

Learning: This is the key phase, in which a significant (or sufficiently large) sample of documents must be selected, and the system must be taught where the required data is located. The AI engine analyzes the texts (NLP technique), both those containing the data to be found and their surrounding context, to find patterns and assign them weights.

Training is done before going into production, although it can always be adjusted later with new samples. As a result, the software will create a knowledge base containing all the pattern analysis.

In this video, I show how the learning phase is carried out using deed samples as an example:

Recognition: In the production phase, the first step is document recognition. At this point, the knowledge base is applied to each document to find patterns that may match each piece of data. In addition to the data itself, the software returns a confidence percentage.

The recognition phase is usually executed autonomously and as an invisible service to users. However, in this video, I show how it would be done manually and how fast it is:

Validation: After recognition, validation is performed on those documents where there is an issue or a field could not be found.

It’s important to remember that the ultimate goal of all these capture solutions is to reduce the total process time. The objective should never be to achieve a certain recognition percentage. In other blog posts, I have analyzed this topic in more detail, since there are still people who focus solely on capture rates.

In this video, I show the results of the previous recognition applied to deeds. In a production solution, the software would only stop at data points with an issue to be resolved by the user, but in this case, I stopped at each data point so you can see the extraction results:

More and more projects are emerging for capturing unstructured documents, and I believe this will be the trend in the coming years. I hope to continue publishing more successful projects showcasing other types of unstructured documents.

]]>
Capturing Documents from a Mobile Device /en/capturing-documents-from-a-mobile-device/ Fri, 10 May 2019 10:30:00 +0000 /?p=341 Continue readingCapturing Documents from a Mobile Device]]>

Have you ever heard the expression “carrying a scanner in your pocket”?

Mobile devices (phones and tablets) offer countless useful features for users, and today I want to highlight the ability to capture documents.

When I talk about capturing a document, I mean obtaining an accurate image of it—not just taking any standard photo. A document image shouldn’t have any kind of background, should be properly rotated, and must be perfectly legible. I can’t count the number of times I’ve come across document repositories full of poorly taken ID photos: rotated, low-quality, and with backgrounds showing tablecloths, tables, fingers, etc. That’s not a document repository, but rather an image repository that can’t be put to much use. For example, you can’t run OCR to extract data, you can’t verify if the document is authentic, and you can’t extract elements like the photo or the signature.

Maybe in our personal lives this isn’t a big deal, but in the business world, information management is fundamental for operations. The higher the quality of a document and the easier it is for a client to send it, the better and faster service the company can provide. This creates a clear competitive advantage, which is why we see examples of companies that allow their clients to send documentation with their mobile devices, or enable their own employees (in branches, remote offices, or even those working on the go) with mobile apps that make it easy to send documents without needing any other device.

To give you an idea of what can be done with a mobile phone, I’ve created the following video, where I show different scenarios in which a document is captured at such a high quality that you can even run OCR and extract data from it.

Do you also use the scanner you carry in your pocket?

]]>
Artificial Intelligence in RPA /en/artificial-intelligence-in-rpa/ Tue, 19 Feb 2019 11:01:00 +0000 /?p=345 Continue readingArtificial Intelligence in RPA]]>

There is a trend in many RPA-related communications to add the words “Artificial Intelligence” (AI) to their messaging, in such a way that one might think a certain RPA technology provider includes Artificial Intelligence in their products.

It’s true that both technologies complement each other well, and from any RPA product, it’s easy to call an external AI service. For example, there are already multiple robots that call Google Cloud Natural Language, IBM Watson, or Microsoft Azure Text Analytics services, among others.

However, what really piqued my curiosity was whether RPA products include any form of artificial intelligence natively (without calling external tools). So, after reading numerous articles and watching many demo videos, I’ve tried to separate the wheat from the chaff and identify what the main RPA providers actually include in their products.

But before anything else, it’s necessary to properly understand some basic terms that are used in this context.

Artificial Intelligence

The Oxford dictionary defines it as: A computer program designed to perform certain operations that are considered to require human intelligence.

One of the most important tools AI uses to solve these problems is Machine Learning, a technique by which systems learn something automatically as they are given information.

Different methods are used for learning: probabilistic, classifiers (support vector machines, nearest neighbor, decision trees…), clustering, regression, etc.

Another term that often appears is Computer Vision, a discipline of AI aimed at enabling a system to understand and classify images as a human would. Machine Learning or other techniques are used for this purpose.

Similarly, Natural Language Processing (NLP) is another AI discipline focused on enabling a system to understand and process human language. Machine Learning techniques are also used to identify structure, language, or specific data. It’s worth noting that only about 10% of a company’s documents contain natural language; most use more constrained language depending on the business type.

Since I can’t analyze all the RPA solutions on the market, I’ve focused on the most prominent ones, starting with Kofax, which I know best.

Kofax

Kofax specializes in automating processes that involve managing information contained in documents. Their solutions include RPA, BPM, multichannel information capture, e-signature, CCM, and BI (Business Intelligence). Their RPA solution includes functionality to classify all types of documents and extract data from them. This allows robots to make decisions based on the content of the documents they handle. Kofax has been using AI in its products for over 15 years but has never advertised it.

Classification uses Machine Learning so that the product is given samples of a document type and automatically generates a knowledge base based on similarities among samples. The same is done for each document type. For example, this function can be used to classify all incoming emails and automatically route them to the appropriate department (accounting, customer service, support, etc.).

Data extraction also uses Machine Learning to determine where the data to be extracted is located in each document and learns over time. Most of the documentation that Kofax handles is business-related, but its most advanced function includes NLP technology to handle natural language, extracting information based on its context. In simpler projects, information is extracted from structured documents like invoices, orders, IDs, contracts, etc., and in more complex ones, from unstructured documents such as mortgages, deeds, meeting minutes, etc.

In addition to Machine Learning, Kofax has a powerful rules engine to complement classification and data extraction.

For screen control, Kofax robots can use Intelligent Screen Automation (ISA), a Computer Vision technique that identifies all objects on the screen and recognizes all visible words (OCR). This allows the robot to navigate the screen based on objects (menus, buttons, text boxes, images, etc.) and the surrounding words, rather than relying on fixed positions. This adds a lot of flexibility in production, as screens don’t always have to be exactly the same, with the same resolution and appearance. The identification of different screen objects was done using Machine Learning, training the software with hundreds of example application screens.

Blue Prism

Blue Prism was one of the pioneers in RPA solutions and, though it’s losing market share, remains one of the sector’s references. You can find hundreds of articles about Blue Prism calling all kinds of external AI services, and the product is ready to communicate with the most well-known ones (Google, Microsoft, etc.).

However, I have not found any reference to actual AI functionality built into the product.

Recently, Blue Prism announced the creation of a new lab dedicated to embedding AI capabilities into its product. https://www.blueprism.com/news/blue-prism-expands-r-d-capabilities-adding-dedicated-ai-labs-and-outlines-roadmap-for-embedded-ai-capabilities

The article highlights that the key will be to include the ability to understand data from documents in any format and use Computer Vision to improve bot design when interacting with environments, following the ideas mentioned earlier.

UiPath

UiPath is one of the fastest-growing companies, to the point that many analysts (e.g., Forrester or Gartner) considered it a market leader in 2018. It’s worth remembering that analysts evaluate not only product functionality, which is what interests me most, but also business parameters like market coverage (industries and geographies), number of references, strategy, business model, etc.

Like other vendors, the product integrates with almost any external AI tool. In addition, a few weeks ago UiPath announced the ability to automate screens using Computer Vision: https://twitter.com/uipath/status/1086231426503106560.

This allows robots to avoid relying on fixed screen positions.Unfortunately, UiPath uses external tools to understand the information in documents, and there do not appear to be plans to change this. It’s the only major vendor for which I could not find any roadmap for adding document understanding features.

Automation Anywhere

Automation Anywhere is the third major player in the RPA sector and likely the leader in the American market. It offers IQ Bot, a technology that enables data extraction from documents using AI techniques: https://www.automationanywhere.com/images/products/IQBotBrochure.pdf

Their messaging is heavily marketing-driven, making it difficult to know exactly which techniques are used, but from the videos I’ve seen, I’d say they use Machine Learning to learn where information is located in documents.

However, the functionality seems rather basic. I haven’t seen examples with complex documents (most of the time, invoices are captured, which has been a solved problem for years), and it doesn’t seem to be able to capture data from unstructured documents (mortgages, contracts, etc.).

I also haven’t seen options to complement learning with design rules (keyword searches, formats, or data relationships) or even create fixed templates. And I’m not clear whether the learning must occur before deployment or if online learning is possible (ideally, both options should be available).

Workfusion

Workfusion’s offering is very similar to Kofax in the sense that, in addition to traditional RPA, the solution includes Machine Learning to capture information from unstructured documents, a workflow manager, and analytics and reporting tools.

Workfusion is a relatively new company, fundamentally rooted in the AI world.

Almost all their public news or videos have a strong marketing component. But these two articles give a good idea of how their main solution (Workfusion SPA) works: 

https://blog.workfusion.com/8-steps-to-supercharging-rpa-7b0982e4c7d3
https://blog.workfusion.com/5-top-questions-email-intake-processing-418ba7905e18

Workfusion’s idea is to use Machine Learning (with different algorithms) to extract information from documents (structured and unstructured) so the robot can perform actions depending on this information. The learning steps are standard (gathering significant samples, user support to show where the data is, and generating the knowledge base).

What the solution lacks, in my opinion, is complementing the AI part with a rules engine to implement the different use cases that commonly arise in these projects.

On the other hand, I’m surprised that, as AI specialists, all their public references talk about capturing very simple documents. For example, Workfusion holds a hackathon among its partners to push its technology to the highest level, and in the two editions held so far, the challenge has been to capture invoice information. As I’ve mentioned before, invoice data capture has been solved for over 15 years. I was hoping to find some more complex use cases.

Others

Most RPA vendors are working on systems that can analyze the work users do to build robots automatically (or more easily). Initially, the idea was fairly simple: record part of the user’s work, and the robot would be built by replaying the recorded tasks. This works well for demos, but it’s not very applicable in production because robots need to learn to handle exceptions, so manual configuration is eventually needed.

That’s why there are companies today trying to apply AI to this design phase, so that after observing a user for days, the software will automatically determine all the decisions the user makes based on the data on screen and implement the robot accordingly. It sounds a bit like science fiction, and currently, this solution works well for creating robots that execute the “happy path” (the task a user performs most frequently), but its designers admit it still has a long way to go before it can correctly recognize exceptions.

Conclusions

Let’s not fool ourselves—traditional RPA products, which only mimic a user’s movements on a PC, don’t require NASA-level technology or highly complex programming. If you want to choose the right one for your business, then you must consider its architecture and scalability, and above all, what gives them added value is the ability to understand the information they process, as this enables the automation of many more processes. AI helps in this area, which is why we’re seeing more and more marketing in this direction.

This is the path RPA vendors are following, although, as we’ve seen, there are significant differences between them today. While Kofax has been capturing document information for many years, others are just getting started, like Workfusion and Automation Anywhere; some haven’t even begun (Blue Prism), and others rely on third-party tools (UiPath), which has many drawbacks—different consultants, different maintenance, different licenses, etc.

The first to market are the ones who have dominated so far, but as we know, in technology, after the initial business development phase comes the competition phase (where we are now, with solutions popping up everywhere), and then comes the domination phase, with one or two solutions as undisputed leaders (which aren’t always the ones who started it all).

We still have some very interesting years ahead in the RPA world.

]]>
Don’t Be Misled by the OCR Percentage /en/dont-be-misled-by-the-ocr-percentage/ Tue, 13 Nov 2018 20:58:00 +0000 /?p=301 Continue readingDon’t Be Misled by the OCR Percentage]]>

I still keep seeing projects where the success of a data capture system is measured based on the OCR accuracy rate. Even in some proof-of-concept tests, clients still tend to compare different solutions according to the extraction percentages obtained.

I suppose this is partly our fault as technology providers, since in the past we focused heavily on this parameter and always tried to improve it as much as possible in our implementations. This year, Kofax published an article on this topic: The Truth About OCR Accuracy, which I want to share with everyone interested in document data capture.

The basic issue is that the capture (or OCR) percentage by itself is not a meaningful business metric. For example, what decision can an executive make if we tell them that one solution has an 80% capture rate and another has 70%? Probably none! How could they understand the impact of either solution on their business, or how would they calculate a possible return on investment? They couldn’t! Most likely, they would request more information to understand the implications of the project for their business.

It would be too simplistic to think that the first solution is better based solely on this figure. What if this first solution has fewer features to facilitate exception handling (data that could not be captured)? Suppose that with the first solution it takes twice as long to resolve each exception. In this scenario, it could actually be faster to handle the 30% of exceptions in the second solution than the 20% in the first. In other words, the second solution could be more effective from a business perspective, offering greater benefits to the client. In fact, the higher the volume of documents to be processed, the greater the benefit compared to the first solution. The article mentioned above also describes techniques that facilitate exception management.

Another factor that can skew our perception is the recognition threshold (the probability limit for accepting data as correct). This is a number (between 1 and 100) that is set manually. Usually, only data with a high recognition threshold (for example, above 80%) is accepted. If the first solution has lowered this threshold significantly (let’s say to 25%), it might get some hits and thus its accuracy rate increases, but you can no longer trust the data it returns, since many of them will be incorrect. For this reason, all information would need to be confirmed manually (everything must be validated because you never know when it will be accurate).

If the second solution has set a higher threshold, it guarantees better quality of the extracted information, but by rejecting more data, its accuracy rate is penalized. The paradox is that both solutions could actually be returning exactly the same data, but the second would appear worse. In both cases, the user must handle the data manually (either by validation or rejection), so once again, the most relevant metric is the time the employee needs to manage exceptions.

In summary, the OCR percentage is an indicative but insufficient metric. What really matters is the total time required to process a document on average, from start to finish. This time depends not only on the OCR percentage but also on the speed at which exceptions are managed. With this information, decisions can actually be made. For example, if a solution allows me to process documents in 70% less time than today’s manual process, I’d very likely be interested in implementing it. If solution A allows me to process 1,000 documents per day and solution B processes 800 documents per day, I can already calculate which one will offer me a better return on investment.

If, despite this, someone is only interested in the OCR percentage, my recommendation would be to simply choose an OCR engine (there are even free ones), not a complete capture solution.

Finally, I would like to highlight that in the modern implementation of these types of solutions, where increasingly complex documents are being handled, the OCR percentage is becoming less and less relevant. Traditionally, when working with more structured documents, rule-based systems were used to capture information. We kept adding more and more rules to improve the percentage, but it became increasingly complicated because each new rule affected the previous ones, so improvements eventually plateaued. Today, projects involve more complex documents such as mortgages, deeds, meeting minutes, etc., and are mostly based on machine learning techniques. In other words, the system is allowed to learn on its own as it processes documentation. No rules are implemented. The challenge with this AI technique is that it needs a lot of samples to learn well. Since there usually aren’t that many examples, projects start with lower recognition rates and the main focus of implementation is on designing effective forms for data correction and entry. The return on investment isn’t as quick, but the cost of manually processing these complex documents is very high, and over time (and more documents) the solution keeps learning and the savings eventually become significant.

]]>
RPA: Buying software or hiring (virtual) workers? /en/rpa-buying-software-or-hiring-virtual-workers/ Sun, 14 Oct 2018 20:43:00 +0000 /?p=298 Continue readingRPA: Buying software or hiring (virtual) workers?]]>

In our offices, just like in life itself, we find all kinds of people: some are friendlier, others better prepared, some communicate more effectively, others work long hours, some have more experience, and so on. We can all agree that diversity enriches companies, and it makes sense to assign each position to the profile best suited for it.

On the other hand, I’ve noticed that the expansion of RPA (Robotic Process Automation) technology in large companies is leading to the creation of robotics departments—sometimes called RPA factories or automation centers—which typically create between 30 and 50 virtual workers (software robots) per year. Some even more.

Following the traditional corporate software purchasing model, these companies have selected a single RPA tool, just as one might choose a single ERP, a single CRM, or a single ECM (Enterprise Content Management) application. As a result, it can—and often does—happen that if this RPA tool isn’t flexible enough, all virtual workers end up being the same and performing very similar tasks, leaving other potentially more profitable processes unautomated.

Some companies already have virtual workers with a good reputation who efficiently automate basic processes, but find themselves stuck when they want to go further. I understand that these companies need to add tools capable of creating virtual workers with more advanced abilities.

For example, virtual workers that can access applications without the need to connect to a PC or desktop. This is the case with Web Automation, where virtual workers can access web pages, web interface applications, Excel files, XML files, or Terminal Server sessions directly in memory, without having to open browsers or connect to other machines. This means huge savings on infrastructure. For instance, using 20 virtual workers for these kinds of tasks translates to saving the infrastructure required for 20 machines or desktops. In fact, I know of two cases with more than 5,000 virtual workers of this kind. I believe this type of robot is essential in any RPA factory.

Another highly valuable feature is the ability to classify documents and extract data from them to make decisions and continue a specific process based on those decisions. This is known as Intelligent Process Automation. It’s true that you can always purchase an external application that your robot can call, but it’s much more advantageous to use a tool that already has this feature integrated. This enables faster and more cost-effective growth by avoiding disparate maintenance, independent updates, and lacking a unified roadmap and architecture. These kinds of robots will be the foundation of the second wave of RPA, which will arrive once the basic processes have been automated. I would also recommend including this type of robot in any RPA factory.

Another point to consider is the way we can invoke our virtual workers. Ideally, they should be callable directly from any application—as a web service, as a scheduled process at regular intervals, or as a batch process (what’s known as “Unattended Automation”)—and they should also be callable manually by selecting them from a list, even with a keyboard shortcut or when an event occurs on our desktop (“Attended Automation”). Most RPA tools offer one or the other way to be invoked, but it’s important that they support as many as possible. Personally, I believe that being able to deploy a robot as a web service has many advantages. Tools that deploy their virtual workers on a server and can be called from any workstation without local installation will always have a competitive edge.

In summary, I think we should ask ourselves whether, with RPA technology, we’re simply buying software or actually hiring workers (virtual ones). In the latter case, I would recommend looking for virtual workers who are capable of taking on multiple job profiles. And if you’re missing a specific profile, the best approach is to “hire” it and add it to your workforce, just as you would with any other employee.

]]>
The True Role of the IT department /en/the-true-role-of-the-it-department/ Tue, 11 Sep 2018 00:01:00 +0000 /?p=286 Continue readingThe True Role of the IT department]]>

At the beginning of the year, while participating in a roundtable on Digital Transformation with CIOs from various sectors, someone asked who should lead projects when a business unit wants to acquire software. The debate was very interesting, with all kinds of arguments. Most stated that in their company, they always select the software themselves. Only a few allow the user to make the selection, with support from the IT department (to ensure integration with other systems, hardware control, and compliance with internal standards).

A couple of months later, during a presentation on the benefits of RPA technology, once again an attendee asked me who should lead these types of projects: the IT department or the user themselves, and a small discussion similar to the previous one began. I shared my experience with this specific technology, for which most of my clients have created an RPA department or team made up mostly of power users who are the ones building the software robots.

But this got me interested in the topic and led me to reflect on the following facts:

  • In my company’s sales methodology, special emphasis is placed on always involving the user, as they are the ones who truly have the problem, understand the software’s functionality best, and know how far it can help them. They can also evaluate how future functionality (the product roadmap) would fit their job in the coming years. On many occasions, we receive a requirements document (RFI or RFP) in which questions about functionality are less important than those about integration, databases, and security.
  • Some of my main clients, large multinational companies with whom I have an ongoing relationship, decided some time ago to decentralize software purchasing, allowing business units to determine which applications they need.
  • A Harvard Business Review survey (2015) on CEOs’ perspectives regarding their IT departments and CIOs revealed that over half of CEOs believe their CIOs don’t know how to effectively apply technology in a changing business, and only 25% think their CIOs perform better than other executives.
  • Department heads now understand much more about technology than they did ten years ago and are perfectly prepared to talk with technology vendors. We also see that more and more CEOs are getting involved in technology matters, and they tend to be the ones who best support digital transformation within their companies.
  • IT departments have spent the last decade carrying out costly and lengthy back-office system implementations such as ERPs, ECMs, CRMs, etc. Besides being seen as a cost center, users tend to think that the only mission of IT is to keep all these systems running day to day and to prevent outages.

With this outlook, I was left wondering about the true role of IT within the company and how I could support them from the perspective of a technology provider. The answer came to me from Jon Mancini, when this summer I took the course “Meeting the Challenge of Digital Transformation” (https://es.linkedin.com/learning/meeting-the-challenge-of-digital-transformation?trk=seo_pp_d_cymbii_title_m015_learning). In one part of the course, Mancini explains that CIOs and their IT departments need to transform themselves because, for many years, they have only focused on one of the letters that gives them meaning: the “T” for Technology, getting involved in continuous application and system installation projects, while completely forgetting about the other letter, the “I” for Information.

As a fundamental premise, the CIO and their IT department are responsible for ensuring that information within a company is available where and when it is needed. From my point of view, this gives a new perspective to the IT department, which should now proactively lead all projects involving information, but not just from the perspective of compatibility, security, or standards, but from the perspective of information itself—which, in the end, is the same perspective as the user’s.

Mancini even believes that a new IT role will emerge, the Information Technician, whose mission will be to understand and care solely about information flows within the company. I believe this is a good way to involve IT in the business and a great way to start an internal digital transformation process.

I suspect that in the coming years, we will see a mix in which many will continue focusing on applications, others will continue empowering their users, and a few will take the lead and start focusing on information flows without being asked by other departments. Surely, the latter will end up gaining a competitive advantage over the rest. What is clear to me is that next time this question arises, the debate will last much longer…

]]>
Spreading One’s Wings /en/spreading-ones-wings/ /en/spreading-ones-wings/#respond Tue, 03 Jul 2018 23:43:00 +0000 /?p=282 Continue readingSpreading One’s Wings]]>

I’m starting this blog with the purpose of having fun, sharing information, and learning. I have spent many years working with different technologies that enable digital transformation within companies, and I would like to share with you my vision on the solutions being implemented, trends, and the significant events I observe in day-to-day work.

I will talk about RPA, which is one of the trendiest technologies lately, and also about Automatic Data Extraction, which, even though it is a very mature market, still has its demand; Electronic Signature, whose adoption keeps growing steadily; Document Composition, another technology in a maturing phase; Mobile Capture, which so far hasn’t seen widespread use except in banking; BPM, and generally about the role that IT departments play today.

Since I understand that my ideas are personal and won’t always be right, I appreciate any comments that might broaden my perspective on any of the published topics. On the other hand, I don’t intend to talk about or compare specific products or brands, but rather focus on technology in general.

Just creating all the infrastructure to write this first post has already been a learning experience for me. As for you, if you find the blog interesting, you can subscribe and share it with your colleagues. And don’t forget that your comments are important.

]]>
/en/spreading-ones-wings/feed/ 0