Building a File Processing Pipeline with Spring Batch and Multi-Step Jobs
Spring Batch remains one of the most reliable frameworks for handling large-scale batch operations in the Java ecosystem. When organisations in Australia process nightly reconciliations, APRA reporting files, or end-of-day ASX settlement feeds, the same pattern keeps surfacing: read a file, transform the records, validate against business rules, and write the cleaned output somewhere useful. Building this as a multi-step job gives you clear boundaries between each phase and makes the whole pipeline easier to monitor, restart, and maintain.
The strength of Spring Batch lies in its chunk-oriented processing model, which handles millions of rows without exhausting memory. Each step in a job can have its own reader, processor, and writer, and the framework takes care of transaction management, checkpointing, and metadata storage. For teams working on integrations with payment gateways or government portals, this approach keeps the codebase predictable and the failure recovery story honest.
This walkthrough takes you through the practical pieces of a file processing pipeline, from project setup to a runnable multi-step job. The examples lean on realistic scenarios you might find in a Sydney-based fintech or a Melbourne logistics firm, including fixed-width files, CSV exports, and the kind of messy input that arrives from partners at 2am AEST.
Core Concepts Behind a Spring Batch Job
A Spring Batch job is a container for steps, and each step is an independent unit of work. The framework relies on a JobRepository to persist execution metadata, a JobLauncher to start jobs, and a JobExplorer to query historical runs. These components work together so that a job that fails halfway through can resume from the last successful checkpoint rather than re-processing the entire file.
Steps come in two flavours: chunk-based and tasklet-based. Chunk steps read, process, and write data in batches of a configurable size, which is the pattern you will use most often for file processing. Tasklet steps suit one-off operations like deleting temporary files, sending a notification email, or invoking a stored procedure at the end of a run.
For Australian compliance work, chunk-based steps are particularly useful because they let you commit transactions in small groups. If a file from a partner contains 50,000 transactions and one row fails validation on record 47,231, the framework skips the bad record and keeps the rest, which is exactly the behaviour APRA-aligned reconciliation systems require.
Setting Up the Project Structure
Start with a Spring Boot project and add the necessary dependencies. You will need spring-boot-starter-batch, spring-boot-starter-data-jpa if you plan to store job metadata in a relational database, and a JDBC driver that matches your target environment. Many Australian teams run on PostgreSQL or Oracle, depending on whether they sit inside a startup or a Big Four bank.
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-batch</artifactId>
</dependency>
Configure your application.yml to point the job repository at a real database during development. In-memory H2 works for quick tests, but production environments in Melbourne and Brisbane data centres typically use persistent storage so the BATCH_JOB_INSTANCE and BATCH_STEP_EXECUTION tables survive restarts. Schedule jobs using a @Scheduled annotation or a quartz trigger, keeping in mind the AEDT to AEST transition that happens each April and October.
When the project needs to scale beyond a single JVM, many teams pair Spring Batch with distributed schedulers. If you are building out a broader platform, you can review how professionals approach enterprise architecture at powerpronet for additional context on production-grade Java systems.
Designing a Multi-Step Job Flow
A multi-step job is the right choice whenever the work cannot be expressed as a single read-process-write cycle. A common Australian banking scenario involves three steps: the first reads a CBA BPay file, the second reconciles it against an internal ledger database, and the third produces a report file that auditors can review. Each step has its own transaction boundary, which prevents a failure in reporting from invalidating the reconciliation work.
Job flows are defined with JobBuilderFactory and StepBuilderFactory. You chain steps using the .next() method for sequential execution, .on() and .from() for conditional branching, and split steps across threads when the work is genuinely independent. For most file processing pipelines, sequential steps are safer because the output of one step often feeds the next.
Consider a job that ingests a fixed-width file containing ASX trade data. Step one parses and normalises the records, step two enriches them with reference data from a database, and step three writes the final output to both a CSV archive and an audit table. This three-step pattern is straightforward to test and easy to explain to a project manager who wants to know what happens at midnight.
Configuring the ItemReader for File Input
The FlatFileItemReader is the workhorse for delimited and fixed-width files. For CSV inputs, configure a DelimitedLineTokenizer with the column names your business uses, and pair it with a BeanWrapperFieldSetMapper to map rows directly onto a domain object. Australian teams often need to handle both comma-separated and pipe-delimited files because partners rarely agree on a single format.
FlatFileItemReader<TradeRecord> reader = new FlatFileItemReader<>();
reader.setResource(new ClassPathResource("input/asx-trades.csv"));
reader.setLineMapper(new DefaultLineMapper<TradeRecord>() {{
setLineTokenizer(new DelimitedLineTokenizer() {{
setNames("tradeId", "settlementDate", "price", "volume");
}});
setFieldSetMapper(new BeanWrapperFieldSetMapper<TradeRecord>() {{
setTargetType(TradeRecord.class);
}});
}});
For fixed-width files, the FixedLengthTokenizer lets you specify column ranges in characters. This is the format used by some legacy mainframe systems in Australian superannuation funds, where record layouts have not changed since the 1990s. Always set a Resource that points to a watched directory rather than a hard-coded path, so the same job can pick up files dropped by upstream systems.
Implementing the ItemProcessor
The processor is where business rules live. Validation, enrichment, filtering, and transformation all happen here. For a payroll file arriving from an outsourced provider, the processor might check that each employee's TFN is present and well-formed, convert the date format from DD/MM/YYYY to ISO standard, and reject rows that do not meet the threshold.
Returning null from a processor tells Spring Batch to skip the item without counting it as a failure. This is useful for filtering out records that are not relevant to the current run, such as zero-value transactions in a reconciliation job. Returning the item unchanged is also valid when the processor only performs side-effects like logging or metrics collection.
Keep processors stateless whenever possible. Australian teams running batch jobs in clustered environments often hit subtle bugs when processors hold state across threads, so favour pure functions and constructor-injected dependencies over mutable fields.
Writing Output with ItemWriter
FlatFileItemWriter handles CSV and fixed-width outputs, while JdbcBatchItemWriter and JpaItemWriter push records into a database. For multi-format outputs, you can use a CompositeItemWriter that fans out to several destinations in a single step. A reconciliation job in a Sydney brokerage might write to both an Oracle ledger and a S3 bucket for archival, satisfying both operational and compliance requirements in one pass.
Configure header callbacks and footer callbacks when downstream consumers expect them. Many Australian partners still rely on header rows containing batch metadata, and missing them can cause whole files to be rejected. The same applies to line aggregates that produce summary records at the end of a file.
Handling Errors, Restarts, and Monitoring
A production file processing pipeline needs a sensible skip policy, a retry policy for transient failures, and a clear restart story. Spring Batch records the last committed chunk in the job repository, so restarting a failed job resumes from that point instead of the beginning. Configure the BATCH_JOB_INSTANCE table to be backed by a database with proper backup procedures, especially when the data must be retained for seven years under Australian tax and superannuation law.
Skip policies tell the framework which exceptions should cause a record to be skipped rather than fail the whole step. Limit the maximum number of skips to avoid silent corruption, and route skipped records to a separate file or SKIPPED table for review. Combine this with a JobListener that posts the summary to Slack or PagerDuty, and your on-call engineer in Adelaide or Perth will know what happened before they finish their morning coffee.