web
You’re offline. This is a read only version of the page.
close
Skip to main content

Announcements

News and Announcements icon
Community site session details

Community site session details

Session Id :
Power Automate
Unanswered

Extract table in PDF

(1) ShareShare
ReportReport
Posted on by 8

Dear Members,

I am new to Power Automate Desktop.

I want to extract this table from this pdf file and convert it to Excel, however I ran into a problem.

When I use the ExtractedPDFtables action, the table only has four columns, not the six that appear in the original PDF. The first two are not extracted.

Could you please explain why and how I could overcome this?

Thanks in advance for your assistance.

I have the same question (0)
  • UshaJyothiKasibhotla Profile Picture
    225 Moderator on at

    Hi @JoaoCord 

    Sorry for the delay...

    I have seen it in unanswered questions.

    Could you please send me the sample pdf so that I can give a try and get back to you.

  • JoaoCord Profile Picture
    8 on at

    Hi,

    I am sending herein attached the pdf file that I would like to get in a excel table.

    The issue is that when i use "ExtractedPDFtables" action from power automate desktop, the table only has four columns, not the six that appear in the original PDF. The first two are not extracted.

    Thanks in advance for your assistance.

    Best regards.

  • CU16071609-1 Profile Picture
    6,255 Moderator on at

    @JoaoCord 

    I think there might be an issue with the PDF table format, resulting in only four columns being displayed instead of six.

    To troubleshoot, I recreated a table(Attached for your reference)  with similar data in Excel, pasted it into a Word document, and then converted the document to a PDF file. Upon extracting data from the recreated PDF file, the correct data was displayed.

    Please see the details of the data tables below for your reference, where I compare your original PDF datatable with the recreated PDF datatable.

     

    Deenuji_0-1708368381638.png

    So I suggested you to correct the table format in the pdf document else extract the data as text using the PAD action "Extract text from pdf" where it will extract the all data but we have to do some workaround to get it formatted.

     

    Extracted data as textL

    Deenuji_1-1708368571883.png

     

    Thanks,

    Deenu

     

     
     
  • JoaoCord Profile Picture
    8 on at

    Hello Deenuji,

    Thanks for your kind help.

    The problem is that I received the pdf file from a professional contact and i should work over it.

    It is a document with multiple pages but in some pages the PAD flow doesn't recognize 6 columns but only 4.

    If i have to do all the steps that you are proposing, besides the time consumption, it will lose all the power that PAD as to offer.

    My goal is to transform the PDF data to excel automatically.

    Thanks again. Joao

  • Agnius Bartninkas Profile Picture
    Most Valuable Professional on at

    The reason for this is likely because the rows in the first two columns are merged, while they are not merged in the other 4 columns. When the row structure is different, PAD does not recognize it as a part of the same table.

     

    The likely only way to work around this would be to extract it as text via Extract text from PDF and then use some regular expressions and other text manipulation actions to retrieve the data you want. But it will still likely be a bit of a pain to do with this sort of a formatting, as it will likely be difficult to identify which rows in the first two columns should correspond to which rows in the other four columns.

  • CU16071609-1 Profile Picture
    6,255 Moderator on at

    @JoaoCord 

    As previously mentioned, the issue lies with the PDF table's formatting, as it can only identify four columns and not the remaining two.

     

    This isn't a critique of PAD's capabilities; rather, it reflects the typical functioning of RPA algorithms. If the table is properly formatted and structured in the PDF, recognition becomes easier. However, if it isn't, resorting to the "Extract text from PDF" action, as mentioned earlier, may be necessary.

     

    Nonetheless, this approach entails significant effort in formatting the table from plain text, where discerning blank columns and values becomes challenging.

     

    Regrettably, this seems to be the only viable solution for your current scenarios.

     

    Thanks,

    Deenu

Under review

Thank you for your reply! To ensure a great experience for everyone, your content is awaiting approval by our Community Managers. Please check back later.

Helpful resources

Quick Links

Season of Sharing Community Challenge Winners!

Congratulations to our community stars!

Kudos to our 2025 Community Spotlight Honorees

Expanding mentorship, skilling, and AI innovation

Congratulations to the June Top 10 Community Leaders!

These are the community rock stars!

Leaderboard > Power Automate

#1
David_MA Profile Picture

David_MA 274 Super User 2026 Season 1

#2
11manish Profile Picture

11manish 175

#3
Haque Profile Picture

Haque 166

Last 30 days Overall leaderboard