How to extract data using AI Extractions
Hi everyone! Sangeetha here, Senior Product Marketing Manager at DataSnipper, with this week's edition of Tips & Tricks.
I used to spend entire engagements retyping numbers out of payroll registers, invoices, and contracts into workpapers. That was the least interesting part of the job, and it was also where one mistyped digit could throw off an entire test. This week I want to talk about AI Extractions, because it takes that whole task off your plate.
AI Extractions pulls fields, tables, and even footnotes out of documents at scale, and every value it pulls links straight back to where it came from in the source file. That link back to source is the part I care about most. You're not just getting a number, you're getting a number you can click on and check against the original PBC document.
Say you're testing payroll for an engagement. You import a batch of payroll reports, click AI Extractions, and point it at the folder. You set up a key like "Payroll Entries" and switch the field type to List, since payroll data is naturally list based rather than a fixed table with the same columns on every page.
From there you add the properties you actually need, things like Name, Hourly Rate, Gross Pay, and 401(k) contributions. Each property gets its own data type. Text works for names and references. Number handles anything with decimals, like pay amounts. Date covers pay periods. Click Run and you get a preview of everything extracted before it ever touches your workbook.
If you're not sure how to set it up or need a good starting point, you can use the Suggest a Schema option in the AI Extractions window. It automatically analyzes your document and recommends a schema, including the data points to extract. It also generates the appropriate data types and extraction descriptions for each field. From there, you can review the suggested schema and make any adjustments to suit your needs
One thing I've picked up from busy season workshops I've run with teams. Sometimes a field pulls from the wrong column. If your 401(k) figure keeps grabbing the wrong number, you can point AI Extractions to exactly where to look, like the Accumulation to Year column, and add a short description so it doesn't confuse similarly named fields. That extra guidance takes thirty seconds and saves you a rework later in the engagement.
If it's a document type you'll see again next quarter, save your setup as a template. That's the difference between rebuilding your extraction logic every time and reusing something you already trust. It also means the next person on the engagement doesn't have to figure out the same field setup from scratch.
Once the preview looks right, export straight to Excel. The data lands in your workbook along with links back to the source documents, so every figure still ties out to your evidence.
Now, the part I always come back to. AI Extractions gets the data out of the document fast, but you're still the one deciding whether it's right. Every extracted value links back to its source, so you can open the original document and confirm it in seconds. That check isn't extra work tacked onto the process. It's what turns an extraction into usable audit evidence instead of just a number sitting on a page. I'd never sign off on a test because the AI produced a figure. I'd sign off because I checked it myself, and the trail back to source made that check fast.
If you want to try this on your next engagement, the full how-to with the field type breakdown and the template steps is here.
I'd love to hear what documents you're testing it against and what tripped you up the first time. Drop a comment and let me know.
I started this series because I've seen how one small tip can completely change the way someone uses DataSnipper. Every week, I'll share something practical you can put to work right away. And if there's a feature or workflow you'd like me to cover next, let me know in the comments. Visit our knowledge base if you'd like to learn more.