web
You’re offline. This is a read only version of the page.
close
Skip to main content

Announcements

News and Announcements icon
Community site session details

Community site session details

Session Id :
Power Platform Community / Forums / Power Automate / Scanned PDF to OCR: Fa...
Power Automate
Answered

Scanned PDF to OCR: Failed to extract text with OCR ---> System.ArgumentException: Parameter is not valid

(0) ShareShare
ReportReport
Posted on by 11

Hello experts,

I'm a beginner at using PAD and I'm trying to get text via OCR out of scanned PDF files. The error message I get - independent of the choice of OCR module is:

Microsoft.Flow.RPA.Desktop.Modules.SDK.ActionException: Failed to extract text with OCR ---> System.ArgumentException: Parameter is not valid. at System.Drawing.Bitmap..ctor(String filename)
at Microsoft.Flow.RPA.Desktop.Modules.OCR.Utilities.Utilities.GetImageForOCR(OCRSource source, SourceScanMode sourceScanMode, Nullable`1 scanRegionX1, Nullable`1 scanRegionY1, Nullable`1 scanRegionX2, Nullable`1 scanRegionY2, IEnumerable`1 imagesToFind, Int32 tolerance, Boolean waitForImage, Boolean timeoutSet, Nullable`1 timeout, Nullable`1 searchRegionImageX1, Nullable`1 searchRegionImageY1, Nullable`1 searchRegionImageX2, Nullable`1 searchRegionImageY2, Action suspendSecureScreen, Action restoreSecureScreen, String imageFilepath, IImageFinder imageFinder)
bei Microsoft.Flow.RPA.Desktop.Modules.OCR.Actions.ExtractTextWithOCRBase.Execute(ActionContext context)
--- Ende der internen Ausnahmestapelüberwachung ---
bei Microsoft.Flow.RPA.Desktop.Modules.OCR.Actions.ExtractTextWithOCRBase.Execute(ActionContext context)
bei Microsoft.Flow.RPA.Desktop.Robin.Engine.Execution.ActionRunner.Run(IActionStatement statement, Dictionary`2 inputArguments, Dictionary`2 outputArguments)

(As I'm working on a German Windows machine I tried to translate at least the beginning of the message back to English but the wording might not match exactly. Unfortunately there seems to be no option to switch languages of PAD.)

As far as I understand the message, PAD fails extracting a getting the image out of the file.

I have tried with a screenshot of the pdf and fed that PNG file to the OCR module, this worked. So could someone tell me what could have gone wrong? Can there be differences in the way a scanned PDF is made up so that PAD might be able or not to extract something?

 

There is another thread on this forum, "How to extract text from PDF using PAD?", that looks like the problem could be the same but I couldn't figure out what the solution post means.

 

Best regards, Joachim

I have the same question (0)
  • Verified answer
    VJR Profile Picture
    7,635 on at

    Hi @JotEss 

     

    If I understood your post correctly I hope my response makes sense to you.

     

    I was not able to directly extract the text from the images inside a PDF. Looks like the PAD action does not support this yet.

     

    I had to first use Extract images from PDF into a folder and then while looping through the images of this folder, passed each of the images to Extract text with OCR.

     

     

     

  • JotEss Profile Picture
    11 on at

    Hi @VJR,

    yes, I guess you understood. I suspected I would have to do something like this and I have already tried "Extract Images from PDF" (or similar wording) (sorry I forgot to mention), but it also resulted in an error. My try to translate:

    The transformation specified is invalid.: Microsoft.Flow.RPA.Desktop.Modules.SDK.ActionException: Error extracting images from „D:\Documents\Scan\Automate\scan.pdf“ ---> System.InvalidCastException: The transformation specified is invalid.
    at System.Linq.Enumerable.<CastIterator>d__97`1.MoveNext() ...

    Original:

    Die angegebene Umwandlung ist ungültig.: Microsoft.Flow.RPA.Desktop.Modules.SDK.ActionException: Fehler beim Extrahieren von Bildern aus „D:\Documents\Scan\Automate\scan.pdf“ ---> System.InvalidCastException: Die angegebene Umwandlung ist ungültig.
    bei System.Linq.Enumerable.<CastIterator>d__97`1.MoveNext()
    bei System.Linq.Buffer`1..ctor(IEnumerable`1 source)
    bei System.Linq.Enumerable.ToArray[TSource](IEnumerable`1 source)
    bei Microsoft.Flow.RPA.Desktop.Modules.PDF.Actions.PDFium.PdfBitmap.GuessPallete(Byte[] indices)
    bei Microsoft.Flow.RPA.Desktop.Modules.PDF.Actions.PDFium.PdfBitmap.ToBitmap()
    bei Microsoft.Flow.RPA.Desktop.Modules.PDF.Actions.ExtractImagesFromPDFAction.Execute(ActionContext context)
    --- Ende der internen Ausnahmestapelüberwachung ---
    bei Microsoft.Flow.RPA.Desktop.Modules.PDF.Actions.ExtractImagesFromPDFAction.Execute(ActionContext context)
    bei Microsoft.Flow.RPA.Desktop.Robin.Engine.Execution.ActionRunner.Run(IActionStatement statement, Dictionary`2 inputArguments, Dictionary`2 outputArguments)

    No clue what is wrong. I found "move next" and thought the algorithm is working on several items in the pdf, so I deleted all pages except the first and deleted an annotation that was contained. But the error persists, even with this cleaned pdf.

    Best regards

    Joachim

  • Verified answer
    JotEss Profile Picture
    11 on at

    I believe I found the cause myself. it was greyscale vs. 1-bit b/w.

    I thought I have to try with a pdf that I am sure I have not yet tampered with, i.e. processed with some editor, before. To make sure it is not some way of edit that makes the pdfs unprocessable. So I happened to grab a scanned page of handwritten notes and it worked (at least extracting the image). And then I noticed, that this was a greyscale scan. Now I tried systematically: Greyscale works, 1-bit black-an-white doesn't!

    This is a pity, because I would prefer 600 dpi b/w to 200 dpi greyscale as it is sharper while only a fraction of the file size.

    Why the h*** did they make this action only work on some sorts of pdf?

    Now I have to feed the images to OCR one by one. Will take me a while as I'm a near complete novice.

  • sdeokule Profile Picture
    31 on at

    Hello VJR

     

    I also need same thing extract data from pdf but in Power automate desktop this action is not available first we need to extract images from pdf. This is unnecessary extra step to follow.

Under review

Thank you for your reply! To ensure a great experience for everyone, your content is awaiting approval by our Community Managers. Please check back later.

Helpful resources

Quick Links

Season of Sharing Community Challenge Winners!

Congratulations to our community stars!

Kudos to our 2025 Community Spotlight Honorees

Expanding mentorship, skilling, and AI innovation

Congratulations to the June Top 10 Community Leaders!

These are the community rock stars!

Leaderboard > Power Automate

#1
David_MA Profile Picture

David_MA 262 Super User 2026 Season 1

#2
11manish Profile Picture

11manish 167

#3
Haque Profile Picture

Haque 154

Last 30 days Overall leaderboard