Prev Next

Cloud / Amazon AppFlow Interview questions

Last updated

1. What is Amazon AppFlow? 2. What are the key components of an AppFlow flow? 3. What are the types of flow triggers in AppFlow? 4. What is a connector in Amazon AppFlow? 5. What is a connection or connector profile in AppFlow? 6. Which sources and destinations does AppFlow support? 7. What is the purpose of field mapping in AppFlow? 8. How do you create a flow in Amazon AppFlow? 9. What are the types of data transfer in scheduled flows? 10. How do you activate, deactivate or delete a flow? 11. List the file formats AppFlow can write to Amazon S3? 12. What is the purpose of filters in AppFlow? 13. What is data masking in AppFlow? 14. How is Amazon AppFlow priced? 15. What is an event-triggered flow in AppFlow? 16. What is the AWS Glue Data Catalog integration in AppFlow? 17. What are custom connectors in Amazon AppFlow? 18. How do you monitor Amazon AppFlow flow runs? 19. How does AppFlow secure data in transit and at rest? 20. Why use AWS PrivateLink with Amazon AppFlow? 21. How does incremental transfer work in AppFlow? 22. How does an event-triggered Salesforce flow work? 23. What is the difference between scheduled and event-triggered flows? 24. What is the difference between Amazon AppFlow and AWS Glue? 25. What is the difference between Amazon AppFlow and AWS DMS? 26. When should you choose AppFlow over a custom Lambda integration? 27. What transformation tasks does AppFlow support? 28. How do you send AppFlow data to Amazon Redshift? 29. How do you transfer data from Amazon S3 into Salesforce with AppFlow? 30. What write operations does AppFlow support for a Salesforce destination? 31. How does AppFlow handle errors during a flow run? 32. What IAM permissions does AppFlow need for an S3 destination? 33. How do you use Amazon EventBridge with AppFlow? 34. How do you automate AppFlow with infrastructure as code? 35. How do S3 partitioning and aggregation work in AppFlow? 36. What happens when a scheduled flow hits the source application's API limits? 37. What are the key quotas of Amazon AppFlow? 38. Explain the execution flow of an AppFlow run? 39. How do you build a Salesforce to data lake pipeline with AppFlow? 40. How can you optimize the cost of scheduled AppFlow flows? 41. How do you troubleshoot a failing AppFlow flow? 42. How do you build a custom connector with the Custom Connector SDK? 43. When would AppFlow not be the right choice? 44. How do you design a multi-account or multi-Region AppFlow setup? 45. How do you protect sensitive data in an AppFlow pipeline? 46. How do you handle source schema changes in AppFlow? 47. How can you trigger downstream processing after an AppFlow run completes? 48. How do you extract data from SAP with AppFlow? 49. How do you design idempotent Salesforce upserts with AppFlow? 50. How do you load data into Snowflake with AppFlow?

1. What is Amazon AppFlow?

Amazon AppFlow is a fully managed integration service that moves data between SaaS applications and AWS services without you writing or hosting integration code. You pick a source, a destination, a trigger and a field mapping, and AppFlow handles authentication, pagination, API throttling and delivery.

A typical use is pulling Salesforce opportunities into Amazon S3 every hour, or pushing records from S3 back into Salesforce.

There are no servers to manage and no upfront fees. You pay per flow run plus the volume of data processed.

Take quiz
AppFlow is best described as:
a container orchestration platform
a managed service that moves data between SaaS apps and AWS services
a CDN for static assets
a relational database engine
How is AppFlow billed?
per provisioned integration server per hour
a fixed yearly license per connector
per flow run plus data processed
per named user seat

2. What are the key components of an AppFlow flow?

A flow is built from a handful of pieces, and interviewers usually expect you to name them in the order you configure them.

  1. Source - the connector and object to read from, such as the Salesforce Account object.
  2. Connection (connector profile) - the stored credentials used to reach the source or destination.
  3. Destination - where data lands, such as an S3 bucket or Snowflake table.
  4. Trigger - on demand, on a schedule or on an event.
  5. Field mapping and tasks - which fields move and how they are transformed.
  6. Filters - optional conditions that limit which records are transferred.
Take quiz
Which component stores the credentials used to reach a SaaS app?
the trigger
the record filter
the connection (connector profile)
the destination bucket policy
Which component decides when a flow starts?
the connector profile
the error bucket
the field mapping
the trigger

3. What are the types of flow triggers in AppFlow?

AppFlow supports three trigger types.

  • Run on demand - you start the flow manually from the console, CLI or API, for ad hoc transfers.
  • Run on event - the flow starts when the source app publishes a supported event, such as a Salesforce change event.
  • Run on schedule - the flow repeats at a fixed interval, as often as once per minute for many sources.

Scheduled flows can do a full transfer or an incremental transfer, and that choice is separate from the trigger type.

Take quiz
Which trigger reacts to a change published by the source application?
run on schedule
run on replication
run on demand
run on event
What is the shortest schedule interval AppFlow supports for scheduled flows?
once per minute (for many sources)
once per hour only
once per five seconds
once per second

4. What is a connector in Amazon AppFlow?

A connector is the component that knows how to talk to one specific application or AWS service. It handles that app's authentication method, its object model, its API limits and how to read or write records.

Salesforce, ServiceNow, Slack, Zendesk, Amazon S3 and Snowflake are all separate connectors. Each connector may act as a source, a destination or both, depending on what the vendor API allows.

If a SaaS app you need is not on the list, you can build a custom connector with the Custom Connector SDK.

Take quiz
What does an AppFlow connector encapsulate?
how to authenticate with and read or write one specific application
the schedule for every flow
a full copy of the source database
a VPN tunnel to your VPC
What can you do if your SaaS app has no built-in connector?
wait for AppFlow to auto-generate one
build a custom connector with the Custom Connector SDK
attach an IAM role to the app
use an S3 lifecycle rule

5. What is a connection or connector profile in AppFlow?

A connection (called a connector profile in the API) stores the details AppFlow needs to sign in to an application, such as OAuth tokens, usernames, passwords or API keys, plus instance URLs.

You create it once and reuse it across many flows. For example, one Salesforce production connection can feed five different flows. Credentials are stored encrypted, using an AWS managed key or your own KMS key.

Rotating a credential means updating the connection, not every flow that uses it.

Take quiz
Why are connections useful?
they replace field mapping
one stored credential set can be reused by many flows
they compress data before transfer
they schedule flows automatically
Where do you update an expired Salesforce OAuth token?
in the S3 bucket policy
in the destination file format
in the connection used by the flows
in every flow's filter

6. Which sources and destinations does AppFlow support?

AppFlow ships with dozens of connectors. Sources include Salesforce, ServiceNow, Slack, Zendesk, Marketo, Google Analytics, Datadog, Dynatrace, SAP OData and Amazon S3.

Destinations include Amazon S3, Amazon Redshift, Snowflake, Salesforce, Amazon EventBridge, Upsolver and Zendesk, among others.

Support is per connector, not universal. Some apps are source-only, and some destinations accept only certain operations, so always check the connector page before you design around one.

A quick way to answer in an interview is to group them: business apps (Salesforce, Zendesk, ServiceNow), marketing and analytics (Marketo, Google Analytics), observability (Datadog, Dynatrace) and AWS or data platforms (S3, Redshift, Snowflake). Custom connectors extend the list when your app is missing.

Take quiz
Which pair is a valid AppFlow source and destination?
Amazon S3 to a local laptop folder
Amazon RDS to Amazon SQS
Salesforce to Amazon S3
DynamoDB Streams to Kafka
Is every connector usable as both source and destination?
No, only S3 can be a destination
Yes, but only in one region
Yes, all connectors are bidirectional
No, support depends on the connector

7. What is the purpose of field mapping in AppFlow?

Field mapping tells AppFlow which source fields go to which destination fields. Without it, the destination would not know where to put each value.

You can map fields one by one, map all fields directly, or upload a CSV of mappings. Mapping tasks also let you rename fields, and this is where transformations such as concatenation, truncation, masking and validation are added.

Mapping is fixed when you save the flow, so a new source field will not appear at the destination until you update the flow or use the map-all option.

Take quiz
What does field mapping define?
how often the flow runs
which IAM role owns the flow
which VPC the flow runs in
which source fields are written to which destination fields
What happens when a new field is added at the source of a flow with explicit mappings?
it is not transferred until you update the mapping
the flow switches to on-demand
it is always transferred automatically
the flow is deleted

8. How do you create a flow in Amazon AppFlow?

From the console the process follows a wizard, and the CLI and API mirror the same steps.

  1. Enter a flow name and optional description, and choose encryption settings.
  2. Choose the source, select or create a connection, and pick the object.
  3. Choose the destination and its options, such as the S3 bucket, prefix and file format.
  4. Choose the trigger: on demand, on event or on schedule.
  5. Add field mappings, optional validations and filters.
  6. Review and create the flow, then activate it or run it.

In code you would use create_flow in the SDK or an AWS::AppFlow::Flow resource in CloudFormation.

Take quiz
Which step comes right after choosing the source and destination in the console wizard?
choosing the trigger
deleting the connection
rotating the KMS key
creating a Glue crawler
Which SDK call creates a flow programmatically?
register_stream
create_flow
start_pipeline
put_flow_rule

9. What are the types of data transfer in scheduled flows?

Scheduled flows offer two transfer modes.

Full transfer Incremental transfer
Sends all records that exist at the time of the run. Sends only records created or changed since the last successful run.
Simple, but re-sends unchanged data every time. Cheaper and faster for large, slowly changing objects.
Good for small reference tables. Needs a timestamp field to detect changes.

If a run fails, an incremental flow picks up from the last successful run so no changes are skipped.

Rule of thumb: start with a full transfer to seed the destination, then switch to incremental for the ongoing loads.

Take quiz
What does an incremental transfer send?
only deleted records
only records added or changed since the last successful run
all records every time
only the first 1,000 records
Which source field is typically required to make incremental transfer possible?
a foreign key to the account table
a primary key of type UUID
a timestamp field marking creation or modification
a binary checksum column

10. How do you activate, deactivate or delete a flow?

Scheduled and event-triggered flows must be activated before they start running. On the flow details page you choose Activate flow, and AppFlow begins running it on its trigger.

Deactivating stops future runs but keeps the flow definition, mappings and connection. You can reactivate it later. Deleting removes the flow completely, and data already written to the destination stays where it is.

On-demand flows have no active state. You simply start a run whenever you need it.

Deactivate a flow before you change its connection or destination, so a half-edited setup does not run. Because runs are billed, deactivating unused scheduled flows also stops the charges.

Take quiz
What does deactivating a scheduled flow do?
deletes the connection
converts it to an event flow
stops future runs but keeps the flow definition
deletes all data already transferred
Which flow type needs no activation?
incremental
event-triggered
scheduled
on-demand

11. List the file formats AppFlow can write to Amazon S3?

When S3 is the destination, AppFlow can write CSV, JSON and Parquet files.

  • CSV - easy to open in spreadsheets and simple tools.
  • JSON - keeps nested structure and is common for API-style data.
  • Parquet - columnar and compressed, the best choice when Athena or Redshift Spectrum will query the data.

The format is set in the destination options, alongside the filename and folder preferences.

If you are unsure, pick Parquet for analytics workloads and JSON when the data is nested or will be consumed by an application. CSV is fine for small extracts that people open in a spreadsheet.

Take quiz
Which S3 output format is columnar and best for analytics engines like Athena?
CSV
plain text
XML
Parquet
Where is the S3 output format chosen?
in the destination settings of the flow
in the source connector profile
in the CloudTrail trail
in the S3 bucket lifecycle rule

12. What is the purpose of filters in AppFlow?

Filters limit which records are transferred, based on conditions you define on source fields. Only records that match are sent to the destination.

For example, you can transfer only Salesforce opportunities where Stage equals Closed Won, or tickets created after a certain date. Filtering at the source reduces the data processed, which lowers cost, and it keeps unwanted or sensitive records out of the destination.

Each source connector supports its own set of filterable fields and operators.

Filters are set per source field with operators such as equals, contains, greater than or between, and several conditions can be combined. Remember that incremental flows still respect these filters on every run.

Take quiz
What do record filters control?
which records are transferred to the destination
how long logs are kept
which region hosts the flow
which IAM users can view the flow
How can filtering reduce AppFlow cost?
it lowers the price per flow run
less data is processed per run
it removes the need for a connection
it disables CloudTrail

13. What is data masking in AppFlow?

Masking is a transformation task that hides part of a field's value before it reaches the destination. You choose the field and a mask character, and AppFlow replaces the characters with it.

A phone number such as 9876543210 could arrive as ******3210 depending on your settings. It is useful for personal data like phone numbers or account numbers that analysts do not need in full.

Masking happens during the flow, so the raw value is never written to the destination.

Take quiz
Where does masking take effect?
only when a bucket is versioned
during the flow, before data is written to the destination
only in CloudWatch logs
only after the data reaches S3
What kind of data is masking mainly meant for?
VPC route tables
IAM policies
sensitive values such as phone or account numbers
binary image files

14. How is Amazon AppFlow priced?

AppFlow pricing has two parts.

  1. Flow runs - a charge per successful run, currently $0.001 per run. A run that checks for new data and finds none still counts.
  2. Data processing - a per-GB charge for the data moved through the service, aggregated across all flows in the account each month. The rate depends on region and on whether the destination is on AWS or uses PrivateLink.

There are no upfront fees. AWS data transfer fees do not apply on top of the data processing charge, though other services you combine with it, like S3 and KMS, bill separately. Check the current pricing page before quoting rates.

Take quiz
Does a scheduled run that finds no new data still count as a flow run?
No, empty runs are free
Only if the run exceeds 1 GB
Yes, it is still billed as a flow run
Only in the first month
Which two factors make up AppFlow billing?
connector licenses and seats
provisioned throughput and storage class
vCPU hours and memory
flow runs and data processed

15. What is an event-triggered flow in AppFlow?

An event-triggered flow starts when the source application publishes a supported event, instead of on a timer. AppFlow receives the event and runs a transfer for the changed data.

A common case is Salesforce change data capture: when an account record is updated, the flow starts and sends the change to S3 or another target. This gives near real-time sync without polling.

Only some connectors expose events, so this trigger is not available for every source.

Because each event starts a run, keep an eye on the flow run count for busy objects. For high-volume sources a short schedule with incremental transfer can be cheaper than reacting to every event.

Take quiz
What starts an event-triggered flow?
a cron expression in the destination
a CloudFormation stack update
a manual click every time
a supported event published by the source application
Which advantage do event triggers have over schedules?
near real-time sync without polling
works with every connector
unlimited data size
no flow run charges

16. What is the AWS Glue Data Catalog integration in AppFlow?

When S3 is the destination, AppFlow can register the transferred data in the AWS Glue Data Catalog. You give it a database name and an optional table prefix, and AppFlow creates or updates a table for each flow.

Once the table exists, you can query the data with Athena, Redshift Spectrum or other catalog-aware tools without running a crawler yourself. Parquet output works best here.

The flow needs an IAM role with permission to create and update tables in the catalog.

Take quiz
What does the Glue Data Catalog integration provide?
an automatically maintained table for the S3 output
a VPN to SaaS apps
a Kafka topic per flow
a replacement for the S3 bucket
What must the flow have to write to the Glue Data Catalog?
a public S3 bucket
an IAM role with catalog permissions
a Lambda function in the same VPC
a NAT gateway

17. What are custom connectors in Amazon AppFlow?

Custom connectors let you connect AppFlow to an application that has no built-in connector, such as an internal API or a niche SaaS product. You build them with the Custom Connector SDK in Java or Python.

The connector runs as an AWS Lambda function that AppFlow calls to fetch metadata and read or write records. Once registered in your account, it shows up like any other connector when you build a flow.

You can also use connectors published by partners through AWS Marketplace.

Take quiz
What do custom connectors run on?
an EMR cluster
an AWS Lambda function
an AppStream fleet
an EC2 Auto Scaling group
Which languages does the Custom Connector SDK support?
COBOL and Fortran
Go and Rust
Java and Python
Only Node.js

18. How do you monitor Amazon AppFlow flow runs?

Start with the run history on the flow details page. It shows each run's status, start time, records processed and the error message for failures.

For automation, AppFlow publishes start and end flow run report events to Amazon EventBridge, and you can route them to Lambda, SNS or Step Functions. Metrics are also emitted to CloudWatch, so you can alarm on failed runs, and CloudTrail records API calls for audit.

A practical setup is one EventBridge rule that filters end reports for failed status and publishes to an SNS topic. That gives the team an email or chat alert without polling the console.

Take quiz
Where can you see the status and error message of a specific run?
the VPC flow logs
the S3 bucket ACL
the run history on the flow details page
the IAM credential report
Which service receives AppFlow run report events for automation?
Amazon Inspector
Amazon Macie
AWS Config rules only
Amazon EventBridge

19. How does AppFlow secure data in transit and at rest?

Data moving between the source, AppFlow and the destination is protected with TLS. At rest, AppFlow encrypts stored items, including connection credentials, with AWS KMS.

By default it uses an AWS managed key. If policy requires it, you can supply a customer managed KMS key when you create the flow or connection, and you then control the key policy and rotation.

Access is governed by IAM, and API activity is logged in CloudTrail. For traffic that must avoid the public internet, add PrivateLink on supported connectors.

In an interview, also mention least-privilege IAM: give the flow creator only AppFlow actions, and scope the S3 bucket policy to the specific bucket and prefix instead of all buckets.

Take quiz
Which key can you supply for at-rest encryption of a flow?
a CloudFront signing key
an ACM public certificate
an SSH key pair
a customer managed KMS key
What protects data in transit between AppFlow and the connected apps?
TLS
plain HTTP with gzip
S3 Transfer Acceleration
AWS Shield Standard

PrivateLink keeps flow traffic on the AWS network instead of the public internet. That reduces exposure to internet-based attacks and satisfies rules that forbid public routes for sensitive data.

It is available for supported connectors such as Salesforce and Snowflake. You enable private connectivity when you create the connection, and the source application must also be set up for it, for example Private Connect on the Salesforce side.

Expect a different data-processing rate for private flows, so compare cost before enabling it everywhere.

Typical reasons are compliance rules for financial or health data, and internal security standards that ban public endpoints. If neither applies, standard TLS over the internet is usually adequate.

21. How does incremental transfer work in AppFlow?

In a scheduled flow with incremental transfer, you choose a timestamp field from the source object, such as LastModifiedDate. On each run AppFlow asks only for records where that field is later than the point reached by the last successful run.

The very first run transfers the data that matches your filters. After a successful run AppFlow records where it stopped, so a failed run does not move the bookmark and the next run retries the same window.

flowchart LR
  A["Scheduled run starts"] --> B["Read last successful run time"]
  B --> C["Query source where timestamp > bookmark"]
  C --> D["Write records to destination"]
  D --> E{Run succeeded?}
  E -- Yes --> F["Move bookmark forward"]
  E -- No --> G["Keep old bookmark and retry next run"]

Pick a timestamp that is updated on every change. If the field only records creation time, edits to old records are missed, which is the most common mistake with incremental flows.

Take quiz
What does AppFlow compare against the bookmark in incremental transfer?
the S3 object size
the selected timestamp field on each record
the connector version
the record's primary key length
What happens to the bookmark when a run fails?
it resets to zero and deletes data
it is copied to CloudTrail
it stays at the last successful point
it moves to the current time

22. How does an event-triggered Salesforce flow work?

You choose a Salesforce change event or platform event as the source, and set the trigger to run on event. AppFlow then listens for those events.

When Salesforce publishes one, for example an update to an Account, AppFlow starts a run and transfers the changed record to the destination. Change data capture must be enabled for the object in Salesforce for change events to be published.

Each event-driven run counts as a flow run for billing, so a busy object can generate many runs in a day.

Because the transfer is driven by events, you avoid polling Salesforce every minute, which also saves Salesforce API calls. Test with a sandbox org first to confirm which events fire for your objects.

Take quiz
What must be enabled in Salesforce for change events to be published?
an S3 Glacier vault
a CloudFront distribution
change data capture for the object
a Redshift cluster
How are event-driven runs billed?
they are free
billed monthly per connector
only the first thousand count
each run counts as a flow run

23. What is the difference between scheduled and event-triggered flows?

Both run automatically, but they start differently and suit different needs.

Scheduled flow Event-triggered flow
Starts on a timer, at most once per minute for many sources. Starts when the source publishes a supported event.
Polls the source, even when nothing changed. Runs only when something changed.
Works with full or incremental transfer. Sends the changed data from the event.
Available for most sources. Available only for connectors that publish events.

Choose events for near real-time sync, and schedules for predictable batch loads.

A useful interview follow-up: a schedule gives predictable cost and load, while events give freshness but a run count that follows source activity.

Take quiz
Which flow type polls the source even if nothing changed?
custom SDK flows only
event-triggered
neither
scheduled
When would you choose an event-triggered flow?
when near real-time sync is needed
when the destination is a local disk
when you want a nightly batch
when the source has no events

24. What is the difference between Amazon AppFlow and AWS Glue?

AppFlow focuses on getting data out of, or into, SaaS apps with no code. Glue is a serverless ETL and catalog service built for heavy transformation of data that already sits in AWS or databases.

Amazon AppFlow AWS Glue
Point-and-click SaaS integration. Spark-based ETL jobs and crawlers.
Light transforms: map, filter, mask, validate. Joins, aggregations and custom logic in code.
Priced per run and per GB processed. Priced per DPU-hour.
Best for ingesting Salesforce, Slack, Zendesk and so on. Best for cleaning and reshaping lake data.

They are often used together: AppFlow lands raw data in S3, then Glue transforms it.

Take quiz
Which service is better suited to heavy Spark-based transformations?
AWS Glue
Amazon SES
Amazon AppFlow
AWS Backup
How are AppFlow and Glue often combined?
AppFlow replaces the Glue catalog
AppFlow lands raw SaaS data in S3, then Glue transforms it
Glue writes to Salesforce, then AppFlow deletes it
Glue runs inside AppFlow

25. What is the difference between Amazon AppFlow and AWS DMS?

They solve different problems. AppFlow integrates SaaS applications through their APIs. AWS Database Migration Service replicates databases, including ongoing change capture from engines like MySQL, Oracle and PostgreSQL.

Amazon AppFlow AWS DMS
Sources are SaaS apps and S3. Sources are database engines.
API-based, subject to the vendor's API limits. Reads database logs for change capture.
Field mapping and simple tasks. Table mappings and transformation rules.

If your data lives in an Aurora or Oracle database, use DMS. If it lives in Salesforce or ServiceNow, use AppFlow.

They can complement each other in one architecture. DMS can replicate the operational database into the lake while AppFlow brings in CRM and support data, and both land in S3 for shared analytics.

Take quiz
Which service replicates data from an Oracle database with change capture?
AWS Amplify
AWS DMS
Amazon AppFlow
Amazon Polly
What does AppFlow read from when pulling Salesforce data?
VPC flow logs
database transaction logs
the vendor's API
EBS snapshots

26. When should you choose AppFlow over a custom Lambda integration?

Choose AppFlow when a supported connector already covers your source and destination and the transformation needs are light. You avoid writing code for OAuth refresh, pagination, throttling retries and schema discovery.

Choose Lambda when there is no connector, when you need complex business logic per record, or when you must call several APIs in sequence during the transfer.

A custom integration also means you own scaling, monitoring and credential rotation, which is the hidden cost that often tips the decision toward AppFlow.

Ask three questions: is there a connector, are the transforms light, and does the team want to own the code? Yes, yes and no points to AppFlow.

Take quiz
Which situation favors a custom Lambda integration?
a ServiceNow-to-S3 nightly load
a scheduled Slack export to S3
complex per-record logic with several API calls
a simple Salesforce-to-S3 copy
What burden does AppFlow remove compared with hand-written code?
S3 bucket naming
IAM policy syntax
DNS management
OAuth refresh, pagination and throttling handling

27. What transformation tasks does AppFlow support?

AppFlow supports light, row-level transformations that you add to the field mapping.

  • Map - link a source field to a destination field.
  • Merge (concatenate) - combine fields, such as first and last name.
  • Truncate - cut a value to a maximum length.
  • Mask - hide characters in sensitive values.
  • Validate - check a field and choose to ignore or stop on failure.
  • Filter - limit which records move.

It does not do joins across objects or aggregations. For those, land the data and process it with Glue, Athena or Redshift.

Validation lets you choose per rule whether to ignore the offending record or stop the flow. Combine it with the error bucket so rejected rows are kept for review.

Take quiz
Which operation is NOT something AppFlow does natively?
masking a field
concatenating two fields
truncating a string
joining two source objects
What does the validate task let you choose?
whether to ignore the record or stop the run on failure
how many retries Lambda gets
which S3 storage class to use
which KMS key alias to rotate

28. How do you send AppFlow data to Amazon Redshift?

AppFlow loads Redshift in two hops. It first writes the data to an intermediate S3 bucket, then issues a COPY into the target table.

You need a Redshift connection with the cluster details and credentials, an IAM role that lets Redshift read the staging bucket, and the target database and table. The table must already exist with columns that fit the mapped fields.

Failed records can be written to an error bucket, depending on your error handling choice.

If the COPY fails, check that the Redshift role can read the staging bucket and that column types line up with the mapped fields. Numeric or date mismatches are the usual cause.

Take quiz
What does AppFlow use between the source and Redshift?
an intermediate S3 staging bucket
a Kinesis stream
a DynamoDB table
an SQS FIFO queue
What must exist in Redshift before the flow runs?
a CloudFront origin
the target table with compatible columns
a Glue crawler
an EC2 bastion

29. How do you transfer data from Amazon S3 into Salesforce with AppFlow?

Use S3 as the source and Salesforce as the destination. AppFlow reads CSV files from the bucket and prefix you choose, maps the columns to Salesforce fields, and writes them with an insert, update, upsert or delete operation.

Keep each CSV file under 125 MB. You can drop several files in the location and AppFlow will process them in a single run, up to the Salesforce export limit of about 500 MB per run.

Column names must line up with the field mapping, and a Salesforce ID or external ID is needed for updates and upserts.

Salesforce field-level security also matters. The integration user must have write access to each field you map, or those columns will be rejected.

Take quiz
What is the per-file size limit for CSV files when S3 is the source and Salesforce is the destination?
1 KB
125 MB
1 TB
5 GB
What do updates and upserts to Salesforce need?
a Glue table
a NAT gateway
an ID or external ID field
a KMS grant to Salesforce

30. What write operations does AppFlow support for a Salesforce destination?

For Salesforce as the destination you choose one of four write operations.

  • Insert - create new records.
  • Update - change existing records, matched on ID.
  • Upsert - update if the record exists, otherwise insert, matched on an external ID field.
  • Delete - remove records identified by ID.

Upsert is the safest choice for repeatable loads because re-running the flow does not create duplicates.

Insert and delete are the riskiest for reruns. Insert duplicates data if you replay a file, and delete cannot be undone from AppFlow, so test on a sandbox first. Update and upsert also require the mapped ID column to be present in every row.

Take quiz
Which operation avoids duplicates when a flow is re-run?
insert
append
upsert
bulk merge without a key
What does update match on?
the flow name
the KMS alias
the S3 object key
the record ID

31. How does AppFlow handle errors during a flow run?

For destinations such as Salesforce, Snowflake and Redshift you set an error handling preference. You either stop the run when a record fails, or ignore the failed record and continue.

Records that fail are written to an S3 error bucket you specify, along with the reason. You can fix the data and re-run, so the bucket doubles as a dead-letter store.

Errors that break the whole run, such as expired credentials, appear in the run history and in EventBridge end-of-run events.

Use stop-on-error for critical loads where partial data is worse than none, and ignore-and-continue for large batches where a few bad rows are acceptable and can be replayed later.

Take quiz
Where do failed records go when you choose to continue on error?
the CloudTrail bucket
a DynamoDB stream
an SQS dead-letter queue
an S3 error bucket
Which choice stops the whole flow at the first bad record?
stop the current flow run
ignore the record and continue
disable validation
switch to on-demand

32. What IAM permissions does AppFlow need for an S3 destination?

Two sets of permissions are involved. The person or role creating the flow needs AppFlow actions such as appflow:CreateFlow, appflow:StartFlow and connection permissions.

The destination bucket needs a policy that lets the AppFlow service principal appflow.amazonaws.com write to it, including s3:PutObject, s3:AbortMultipartUpload, s3:GetBucketLocation and the multipart listing actions.

If you use a customer managed KMS key, its key policy must also allow AppFlow to use it.

A common mistake is fixing the IAM policy of the user while leaving the bucket policy unchanged. The flow writes as the service, so the bucket policy is what decides whether the write succeeds. Scope it to the exact bucket and prefix.

Take quiz
Which principal must the destination bucket policy allow?
appflow.amazonaws.com
s3.amazonaws.com only
lambda.amazonaws.com
glue.amazonaws.com
What extra grant is needed if a customer managed KMS key is used?
an ECR pull permission
a key policy that allows AppFlow to use the key
a Route 53 record
a public bucket ACL

33. How do you use Amazon EventBridge with AppFlow?

EventBridge plays two roles. It can be a destination, where AppFlow sends each record from a SaaS source as an event on an event bus, so rules can fan them out to Lambda, SQS or other targets.

It is also where AppFlow publishes flow run reports, with start and end events that you can match with a rule to send alerts or start the next step of a pipeline.

Events sent as a destination are limited to 256 KB. For larger events AppFlow publishes a summary that points to the full data in S3.

Use the destination role for streaming SaaS changes into an event-driven design, and the run report role for operational alerts. They are independent and can be used in the same account.

Take quiz
What size limit applies to events AppFlow sends to EventBridge?
256 GB
256 KB
10 bytes
1 GB
What can an EventBridge rule do with an end-of-run report?
resize the S3 bucket
change the source schema
start an alert or the next pipeline step
rotate the connection secret

34. How do you automate AppFlow with infrastructure as code?

CloudFormation has AWS::AppFlow::ConnectorProfile and AWS::AppFlow::Flow resources, and Terraform has matching AppFlow resources. Defining flows this way makes them repeatable across accounts and reviewable in pull requests.

A minimal on-demand Salesforce-to-S3 flow looks like this:

Resources:
  AccountToS3:
    Type: AWS::AppFlow::Flow
    Properties:
      FlowName: sf-account-to-s3
      TriggerConfig:
        TriggerType: OnDemand
      SourceFlowConfig:
        ConnectorType: Salesforce
        ConnectorProfileName: sf-prod
        SourceConnectorProperties:
          Salesforce:
            Object: Account
      DestinationFlowConfigList:
        - ConnectorType: S3
          DestinationConnectorProperties:
            S3:
              BucketName: my-lake-bucket
      Tasks:
        - TaskType: Map_all
          SourceFields: []
          ConnectorOperator:
            Salesforce: NO_OP
          TaskProperties:
            - Key: EXCLUDE_SOURCE_FIELDS_LIST
              Value: '[]'

Keep secrets out of templates by creating the connection with values from Secrets Manager or a secure parameter.

Treat the snippet as a starting point. Add a schedule trigger, filters and an S3 output format as needed, and validate the template in a test account before promoting it.

Take quiz
Which CloudFormation resource type defines an AppFlow flow?
AWS::Glue::Trigger
AWS::Events::Connection
AWS::AppFlow::Flow
AWS::Lambda::EventSourceMapping
Which trigger type does the sample template use?
Continuous
OnEvent
Scheduled every minute
OnDemand

35. How do S3 partitioning and aggregation work in AppFlow?

For an S3 destination you can control how output files are laid out.

  • Folder structure - partition output by year, month, day or hour so query engines can skip irrelevant folders.
  • Aggregation - either write records as they come, or aggregate all records from a run into a single file, with an optional target file size.
  • Filename options - add a timestamp so repeated runs do not overwrite each other.

Partitioned Parquet plus the Glue Data Catalog is the usual setup for cheap Athena queries.

Aggregating into one file per run reduces small-file problems for Athena, while hourly folders help when you need fine-grained time filtering. Pick the granularity that matches how people query the data.

Take quiz
Why partition S3 output by date?
it removes the need for a bucket
it encrypts each file separately
it makes files smaller than 1 KB
query engines can skip folders and scan less data
What is the purpose of adding a timestamp to filenames?
to avoid overwriting files from earlier runs
to speed up TLS
to skip the mapping step
to lower KMS cost

36. What happens when a scheduled flow hits the source application's API limits?

AppFlow calls the source through its public API, so it draws from the same quota as your other integrations. On Salesforce, for instance, flow calls count against the org's API limits.

When the limit is hit, calls are throttled or rejected, and the run can fail or slow down. Fix it by lowering the schedule frequency, using incremental transfer, adding filters, or letting the connector use a bulk API where available.

Also note the connector-specific ceilings, such as how often some sources can be scheduled, before you set an aggressive interval.

Monitor API usage in the source app itself, since AppFlow is only one of the consumers. Coordinate schedules so several flows do not start on the same minute.

Take quiz
How can you reduce API pressure on a source app?
use incremental transfer and lower the schedule frequency
disable filters
add more S3 buckets
increase the schedule to every second
What can a Salesforce flow's calls consume?
only KMS requests
API limits of the Salesforce org
the org's S3 credits
AWS Shield tokens

37. What are the key quotas of Amazon AppFlow?

These defaults matter most when you design at scale. Some can be raised through Service Quotas and some are fixed.

Quota Default
Flows per account (per Region) 1,000
Flow runs per month 10,000,000
Concurrent flow runs 1,000
Connector profiles 100
Maximum time for one flow run 48 hours
Fastest schedule for many sources 1 run per minute

Individual connectors add their own limits, such as Salesforce import size per run, so read the connector notes too.

Request increases through Service Quotas for the adjustable ones, such as flows per account and monthly runs, before a launch rather than after a failure. Fixed limits like the per-run duration need design changes instead.

Take quiz
What is the default maximum time one flow run can take?
15 minutes
48 hours
5 minutes
7 days
How many flows can one account have per Region by default?
50,000
10
1,000
unlimited

38. Explain the execution flow of an AppFlow run?

A run follows the same path regardless of trigger. AppFlow checks the connection, reads from the source in pages, applies mapping tasks, and writes to the destination.

flowchart TD
  A["Trigger fires"] --> B["Validate connection and permissions"]
  B --> C["Read source records in pages"]
  C --> D["Apply filters and mapping tasks"]
  D --> E["Write to destination"]
  E --> F{Any failures?}
  F -- Records failed --> G["Write to error bucket per preference"]
  F -- None --> H["Publish run report to EventBridge"]
  G --> H

For incremental flows the bookmark is moved only after a successful run. The end report includes records processed and the final status.

Notice that errors can occur at several points. A failure before reading, such as a bad connection, stops the run at once. A failure while writing may affect only some records, depending on your error preference.

Each run also produces a report with the number of records processed, which is what monitoring tools consume.

Take quiz
When is the incremental bookmark moved?
on every retry
when the flow is created
after a successful run
before reading the source
Which step comes immediately after reading source records?
rotating credentials
creating the bucket
publishing to CloudTrail
applying filters and mapping tasks

39. How do you build a Salesforce to data lake pipeline with AppFlow?

Land raw Salesforce data in S3, register it in the catalog, and query it with Athena.

flowchart LR
  SF[Salesforce] --> AF["AppFlow scheduled flow"]
  AF --> S3["S3 partitioned Parquet"]
  S3 --> GC["Glue Data Catalog table"]
  GC --> AT["Athena or Redshift Spectrum"]
  1. Create a Salesforce connection and a flow with incremental transfer on LastModifiedDate.
  2. Write Parquet to S3, partitioned by day.
  3. Enable the Glue Data Catalog option so a table is kept up to date.
  4. Query with Athena, and use an upsert-style view or job to keep the latest version of each record.

Upserts to the lake are not automatic, since S3 files are immutable. Either keep every version and pick the latest with a query, or run a Glue or Athena job that merges changes into a curated table.

Add a second flow for related objects such as Contact or Opportunity so the lake can be joined later.

Take quiz
Which S3 format suits this lake pipeline best?
plain log text
zip archives
single large CSV
partitioned Parquet
What do you use to detect changed Salesforce records?
incremental transfer on LastModifiedDate
full transfer every minute
S3 versioning
an SQS delay queue

40. How can you optimize the cost of scheduled AppFlow flows?

Cost has two drivers, so tune both.

  • Fewer runs - every run is billed, even an empty one, so use the longest interval the business accepts.
  • Incremental transfer - avoids re-processing unchanged records.
  • Filters and field selection - transfer only the rows and columns you need.
  • Retire idle flows - deactivate flows nobody uses.
  • Prefer Parquet - smaller files reduce downstream storage and query cost.

Compare event triggers against schedules for spiky data, because a busy source can generate more runs than an hourly schedule would.

Do the math once: an every-minute flow makes about 43,200 runs per month, versus 720 for an hourly flow. At the per-run price that difference is visible, before any data processing charges.

Take quiz
Which change reduces the number of billed runs?
using a longer schedule interval
switching to JSON
enabling PrivateLink
adding more filters only
Why can event triggers cost more than an hourly schedule for busy sources?
events disable KMS
each event-driven run is billed
events use a different currency
events force full transfer

41. How do you troubleshoot a failing AppFlow flow?

Work from the outside in, using the error in the run history as your first clue.

  1. Credentials - expired OAuth tokens or changed passwords in the connection are the most common cause.
  2. Source limits - check for API quota errors on the source app.
  3. Mapping and data - a renamed or deleted field, or a value that fails validation.
  4. Destination permissions - the S3 bucket policy, the KMS key policy or the Redshift role.
  5. Error bucket - read the failed records and reasons written there.

Then create an EventBridge rule on the end-of-run report so future failures alert you immediately.

Reproduce with a single on-demand run after each fix, rather than waiting for the next schedule. If the error message is vague, check CloudTrail for the API call that was denied.

Take quiz
What is the most common cause of sudden flow failures?
too many S3 folders
expired credentials in the connection
a full CloudTrail log
long bucket names
Where can you read the reasons individual records were rejected?
the CloudFront log
the Route 53 zone
the S3 error bucket
the IAM access advisor

42. How do you build a custom connector with the Custom Connector SDK?

You implement three handlers in Java or Python and deploy them as an AWS Lambda function.

  • ConfigurationHandler - validates settings and credentials.
  • MetadataHandler - lists the objects and fields the connector exposes.
  • RecordHandler - reads records for queries and writes records for a destination.

After deploying the Lambda, you register the connector in your account so it appears in the console, then create connections and flows with it. Test with a small object first, and handle pagination and rate limits inside the handler because AppFlow relies on it.

Keep the Lambda timeout and memory realistic for the size of the pages you return. If you plan to share the connector, package it for AWS Marketplace or share it across accounts through the registration process.

Take quiz
Which handler lists the objects and fields a connector exposes?
RecordHandler
BucketHandler
MetadataHandler
ConfigurationHandler
How is a custom connector deployed?
as a Glue crawler
as an S3 static site
as an ECS task definition
as an AWS Lambda function

43. When would AppFlow not be the right choice?

AppFlow is a poor fit when the job needs more than light, row-level transfers.

  • Heavy transformation - joins, aggregations or custom logic belong in Glue, EMR or SQL.
  • Sub-minute latency - schedules bottom out at one minute, so true streaming needs Kinesis or MSK.
  • Database replication - use DMS for change capture from database engines.
  • No API or connector - a system with no usable API needs a custom connector or another tool.
  • Very API-limited sources - the source quota may cap what any tool can pull.

A good answer states the alternative too: Glue for transformation, Kinesis for streaming, DMS for databases, and a custom connector or Lambda when the app has no supported connector.

Take quiz
Which requirement points away from AppFlow?
masking a phone number field
a nightly Salesforce export
loading CSV files into Salesforce
streaming data with sub-second latency
Which service is better for ongoing database change capture?
AWS DMS
AppFlow
Amazon Rekognition
AWS Backup

44. How do you design a multi-account or multi-Region AppFlow setup?

AppFlow flows and connections are Regional, so create them in the Region where you want to run them. A flow in one Region cannot be managed from another.

For a central data lake in another account, write to an S3 bucket in that account and grant the AppFlow service principal access in the bucket policy, plus KMS permission if the bucket is encrypted with a customer managed key. Where data must move between Regions after landing, use S3 replication.

Deploy the same flow templates to each account with CloudFormation StackSets so definitions stay consistent.

Remember that connections must also be created in each Region and account where flows run, so credentials should be managed centrally, for example through a shared secrets process, to avoid drift.

Take quiz
Are AppFlow flows Regional or global?
Regional
account-wide across all Regions
global by default
tied to one Availability Zone
How can you move landed data to another Region?
by editing the connector profile
S3 replication
by enabling a global flow
by using IAM Identity Center

45. How do you protect sensitive data in an AppFlow pipeline?

Layer several controls rather than relying on one.

  • Mask or drop sensitive fields in the mapping so raw values never land.
  • Encrypt with a customer managed KMS key for flows, connections and the S3 bucket.
  • Keep traffic private with PrivateLink on supported connectors.
  • Restrict IAM so only specific roles can create or start flows.
  • Scan S3 with Amazon Macie to detect stray personal data.
  • Audit with CloudTrail.

Do not forget the downstream copy. Even if the flow is secure, an analyst export or a public bucket undoes the protection, so apply the same controls to the whole path through S3, Glue and Athena.

Review the field list with the data owner before the first run, since removing a column later does not remove data already landed.

Take quiz
Which service can scan the S3 destination for sensitive data?
Amazon Lex
Amazon Macie
AWS Shield
AWS Snowball
What is the safest way to keep a sensitive field out of the destination?
rename the flow
increase the schedule
leave it out of the mapping or mask it
use JSON instead of CSV

46. How do you handle source schema changes in AppFlow?

Mapped fields are fixed when you save a flow. If a field is renamed or removed at the source, the flow can start failing until you update the mapping.

If a field is added, it is skipped when you mapped fields explicitly. Using the option that maps all fields directly picks up new fields, but the destination schema must be able to accept them, which is easy in S3 and harder in Redshift or Snowflake.

Add an alarm on failed runs and review source release notes so you catch changes before they hit production.

For a controlled process, keep flow definitions in version control, and treat a source schema change as a small release: update the template, test in a sandbox account and then promote.

Take quiz
What happens to a newly added source field when only explicit mappings exist?
it is masked automatically
the flow is deleted
it is not transferred
it is transferred and the flow is renamed
Which destination is most tolerant of new fields?
a Salesforce required field
a Snowflake table with fixed columns
a strictly typed Redshift table
Amazon S3 files

47. How can you trigger downstream processing after an AppFlow run completes?

There are two clean options.

  1. EventBridge rule on the end flow run report. Filter on the flow name and a success status, then target a Lambda function or Step Functions state machine.
  2. S3 event notification on the destination bucket. New objects trigger Lambda, SQS or SNS.
flowchart LR
  AF["AppFlow run ends"] --> EB["EventBridge end report"]
  EB --> R{Rule matches flow and status}
  R --> SF["Step Functions or Lambda"]
  SF --> N["Glue job or notification"]

The EventBridge route tells you the run finished, including failures. The S3 route only fires when files actually appear.

Prefer EventBridge when downstream work must wait for a complete run, because S3 events fire per file and a multi-file run would trigger the next step several times.

Include the flow name in the rule pattern so one rule does not start work for unrelated flows.

Take quiz
Which signal tells you a run ended, including failures?
an S3 object-created event only
a CloudFront access log
a Route 53 health check
the EventBridge end flow run report
What can an S3 event notification target after new files arrive?
Lambda, SQS or SNS
a VPC peering link
an IAM user
a Route 53 zone

48. How do you extract data from SAP with AppFlow?

AppFlow's SAP OData connector reads from OData services exposed by your SAP system. You choose the service and the entity to extract, then map and filter fields like any other source.

You supply the application host, service path and authentication, using basic or OAuth. If the SAP system is not reachable from the internet, set up private connectivity through a VPC endpoint service so traffic stays on your network.

Start with small entities and test pagination and filters before scheduling large extractions, because SAP performance limits apply.

Scope filters carefully. Pulling an entire large entity through OData can be slow and load the SAP system, so filter by date or company code and schedule outside peak business hours.

Take quiz
What does the SAP OData connector read from?
OData services exposed by SAP
SAP HANA redo logs
a Kafka topic
an EBS snapshot
How can you avoid the public internet for a private SAP system?
make SAP public
use private connectivity through a VPC endpoint service
disable authentication
use a browser plugin

49. How do you design idempotent Salesforce upserts with AppFlow?

Idempotent means running the flow twice leaves Salesforce in the same state. Use the upsert operation with an external ID field that uniquely identifies each record in your source data.

  1. Create an external ID field on the Salesforce object and mark it unique.
  2. Include that column in the S3 CSV or source data and map it.
  3. Choose upsert and select that field as the ID.
  4. Set error handling to continue and review the S3 error bucket after each run.

A retry after a partial failure then updates existing rows instead of creating duplicates.

Avoid using the Salesforce record ID as the key when the source system does not know it. A separate external ID, such as the ERP customer number, lets both systems refer to the same record.

Take quiz
Which operation makes repeated runs safe?
insert without a key
upsert on a unique external ID
delete then insert
update on the flow name
What should the external ID field be?
random per run
the flow ARN
unique per record
empty for new rows

50. How do you load data into Snowflake with AppFlow?

AppFlow writes to Snowflake through an S3 staging bucket and a Snowflake stage. It puts the files in S3 and then loads them into the target table.

You configure a Snowflake connection with the account, warehouse, stage and storage integration details. Then you choose the database, schema and table in the destination settings.

Snowflake supports private connectivity with PrivateLink, and failed records can go to an error bucket, so both security and troubleshooting are covered.

Set up the storage integration so Snowflake can read the staging bucket, and grant the AppFlow role access to the same bucket. Most first-time failures come from a missing bucket permission on one side.

Make sure the Snowflake table exists with compatible column types before the first run.

Take quiz
What does AppFlow use to stage data for Snowflake?
an RDS read replica
an SNS topic
an S3 bucket and a Snowflake stage
a Kinesis shard
Where do you choose the target table for Snowflake?
in the source filter
in the CloudWatch dashboard
in the S3 lifecycle policy
in the destination settings of the flow
«
»

Comments & Discussions