top of page

Turnaround Production Systems

  • Jul 16
  • 8 min read

ABSTRACT

Oil, gas, and chemicals are often produced at scale in continuous-processing facilities that operate around the clock for years. These plants shut down equipment only every 2-5 years to perform scheduled maintenance and upgrades, known as shutdowns, turnarounds, or outages (STOs). These events occur in high-volume, low-margin commodity industries, where aging assets, shrinking refining capacity, and volatile crack spreads amplify the financial impact of downtime. With more than 65% of turnarounds exceeding planned cost or schedule, the ability to execute predictable, efficient outages has become a strategic differentiator. When a refinery goes offline, planned or unplanned, the economic stakes are immediate: lost production, elevated market prices, and cascading impacts on reliability. Turnarounds require thousands of simultaneous tasks performed by multiple contractors under hazardous conditions, in which queueing delays, permitting constraints, and bottlenecks often dominate actual cycle times. This paper demonstrates how operations science and Production System Optimization (PSO) provide a structured framework for optimizing turnaround execution, offering industry professionals a proven method to improve outcomes with confidence.

 

Keywords: Shutdowns, Turnarounds, Outages


INTRODUCTION

According to industry data, more than 65% of turnarounds exceed planned cost or schedule [1]. In a high-volume, low-margin industry, this level of underperformance carries significant economic consequences. Driven by macroeconomic pressures, regulatory constraints, and volatile crack spreads, many U.S. refineries have shuttered operations in recent years. What remains is an aging fleet of large, highly complex facilities that are critical to the supply of fuels and petrochemicals. The average U.S. refinery is now over 50 years old, and rather than building new capacity, operators have relied on incremental expansions and retrofits to meet demand [2]. U.S. refining capacity declined significantly following the COVID-19 period, falling by approximately 1.1 million barrels per day, roughly 5% of total capacity, in a short period due to refinery closures and conversions [3].

 

Turnarounds are therefore not routine maintenance events; they are high-stakes operational and economic inflection points. When a refinery goes offline, even for planned maintenance, the market often responds immediately with price increases and tighter supply. Every additional hour of downtime carries a disproportionate cost, elevating the importance of predictable, efficient execution.

 

A turnaround is not a single project but a complex program comprising hundreds or thousands of interdependent work scopes executed across multiple process units. These activities are performed by shared resources, including craft labor, inspectors, permitting authorities, and specialized contractors, operating within constrained and often hazardous environments. Work is executed in parallel across congested workfaces, with coordination requirements spanning disciplines, contractors, and physical locations. While individual tasks may appear routine, execution conditions vary significantly depending on unit-specific hazards, permitting requirements, and access constraints. For example, a valve replacement in a utility system may be straightforward. At the same time, the same task in a high-hazard unit can require substantially more time due to additional controls, PPE, and inspection hold points.

 

Traditional planning and scheduling approaches treat turnarounds as deterministic sequences of tasks with defined durations. In practice, execution behaves as a stochastic production system governed by variability, resource constraints, and queueing dynamics. Across many workflows, queue times, not raw process times, dominate total cycle time, as work is frequently delayed by permitting, scaffold access, inspection hold points, and equipment isolation constraints. As a result, schedule performance is driven less by the speed of individual tasks and more by the efficiency of flow through the system.

 

This distinction fundamentally changes how turnarounds must be planned and executed. When production rates exceed the service capacity of constrained resources, queues, or the time spent waiting on other work, they expand nonlinearly, creating bottlenecks that drive disproportionate schedule delays. These bottlenecks are often not the primary craft activities but shared services. Without explicitly modeling these dynamics, even well-developed schedules tend to drift as variability compounds across the system, increasing the risk of delays that industry professionals seek to avoid.

 

Production System Optimization (PSO) provides a structured framework for addressing these challenges by modeling turnarounds as integrated production systems. By explicitly accounting for flow variability, resource constraints, and workflow dependencies, PSO directly tackles industry concerns about schedule overruns and unpredictability, enabling the identification of bottlenecks, the quantification of queueing effects, and the determination of Last Responsible Moment (LRM) decision points. This approach shifts planning from static schedules to dynamic, flow-based execution strategies, aligning with industry needs for more reliable turnaround outcomes.

 

Equally important, PSO enables organizations to manage uncertainty during execution. Emergent work, driven by previously unknown equipment conditions or higher-than-expected degradation, can be evaluated in real time within the context of the overall system. Rather than relying on static contingency buffers, operators can make informed decisions about work sequencing, resource allocation, and capacity adjustments to maintain schedule control and foster confidence in handling unforeseen challenges.

 

By reframing turnarounds as dynamic production systems rather than static project schedules, organizations can improve predictability, reduce schedule overruns, and accelerate safe return to service. In an environment where downtime directly impacts market supply and profitability, the ability to control flow is not just an operational advantage; it is a strategic necessity.

Turnaround planning


Turnaround programs consist of both routine maintenance activities and capital work scopes, including piping, pressure vessel, and storage tank repairs, replacements, and alterations. Due to long-lead materials and fabrication requirements, planning efforts often begin up to two years before the scheduled outage. During this phase, inspectors, engineers, and planners develop discipline-specific work packages that define scope, execution methodology, materials, and quality requirements. These packages are then sequenced for execution based on unit location, fluid system, and equipment criticality.

 

While this structured planning process establishes the foundation for execution, it is inherently limited by uncertainty in actual field conditions and by its reliance on deterministic assumptions. The prevailing objective during planning is to maximize prefabrication and off-site work to increase on-site throughput and reduce outage duration. Prefabrication shifts labor to controlled environments, improves quality, and reduces field welding and fit-up time. However, it also introduces coordination challenges during execution.


A common issue is the “matching problem,” where preassembled piping spools or components are staged in laydown yards well in advance of installation. During execution, craft labor expends significant time searching for, retrieving, and repositioning materials, creating hidden inefficiencies that erode the gains achieved through prefabrication. Excessive laydown area congestion further compounds this issue, increasing handling time, risk of damage, and overall cycle time.

 

A flow-based approach to planning addresses this constraint by aligning material delivery with installation readiness. Delivering piping spools and components to the workface at or near the Last Responsible Moment (LRM), rather than staging them months in advance, reduces congestion, minimizes handling, preserves material integrity, and shortens installation cycle times. This approach requires tighter integration among procurement, fabrication, and field execution, but it significantly improves overall system flow.

The oil and gas industry has made substantial investments in improving turnaround predictability through planning standardization, centralized execution teams, and the adoption of advanced technologies. Industry-leading organizations leverage 3D modeling, digital twin environments, and LiDAR-based reality capture to enhance scope definition, identify potential clashes, and improve workface planning. These tools reduce uncertainty before execution but do not eliminate the variability introduced during the event itself. As a result, even the most advanced planning tools must be complemented by execution strategies that explicitly manage variability and flow during the event.


Flow Determines Schedule: The Role of Production Rates


Standard work processes define how tasks are executed and help transfer tribal knowledge, but they fail to capture the full system when variability and queueing effects are ignored. A pipe field welding task may require only one hour of touch time; however, when accounting for permit issuance, scaffold inspection, and required quality hold points for fit-up, root, and cap, as much as half a workday can be consumed in waiting. On a recent shutdown, waiting accounted for 60–80% of total cycle time across welding workflows.

 

When production rates exceed the service capacity of permitting or inspection processes, queues expand nonlinearly, driving schedule slippage that is disproportionate to task durations. Bottlenecks in turnaround execution are therefore most often found at inspection hold points, scaffold access constraints, or equipment isolation steps, and not within the craft execution itself. Flow efficiency, not individual task productivity, determines overall schedule performance.

 

Many turnaround activities are repetitive and have been executed before, leading planners to reuse prior execution plans. However, no two tasks are identical. Variations in workface congestion, procedural changes, safety requirements, or operating conditions introduce significant variability into execution. For example, depositing defect-free welds on previously in-service piping is inherently more difficult due to contaminant release during welding, which increases the risk of rework and extends cycle times. Applying standard cycle times developed for new construction to these conditions introduces systematic error into schedule forecasts.

 

Accurate scheduling requires modeling production systems as they will actually perform, including variability, resource constraints, and all workflow handoffs. This approach enables identification of the Last Responsible Moment (LRM), the latest point at which a task or decision must occur to avoid delaying downstream work. Planning against LRM dates supports just-in-time execution, reduces material congestion, and aligns work sequencing with actual system capacity.

 

Traditional scheduling methods treat turnaround execution as a deterministic sequence of tasks. In practice, turnarounds behave as stochastic production systems governed by variability, constraints, and queueing dynamics. Without explicitly modeling these factors, even well-developed plans drift from baseline due to compounded inefficiencies. A production-based perspective reframes the turnaround as a dynamic flow system rather than a static schedule, enabling more predictable, efficient, and controlled execution.


Managing Uncertainty During Execution


Most equipment in a turnaround has not been internally inspected since the previous outage. While external inspections are performed prior to the event with systems in service, they provide only limited visibility into internal degradation mechanisms. As a result, actual field conditions frequently diverge from pre-turnaround assumptions. Corrosion rates may exceed projections, previously undetected damage may be uncovered, or components may be found in a condition that necessitates immediate repair or replacement to ensure safe and reliable operation through the next cycle.

 

This inherent uncertainty is a defining characteristic of turnaround execution. To account for it, contingency is typically embedded within both the schedule and budget. However, contingency is often treated as a static buffer rather than a dynamic management tool. When multiple emergent work scopes are introduced simultaneously, static contingency can be quickly consumed, resulting in cascading delays and loss of control over the critical path.

 

A production system perspective provides a more effective framework for managing this uncertainty. Rather than passively absorbing variability, a modeled production system enables real-time evaluation of emergent work. By incorporating current system conditions, such as crew availability, permitting capacity, inspection resources, and workface access constraints, decision makers can assess multiple execution scenarios and determine the optimal path forward.

 

For example, when unexpected vessel repairs are identified, the decision is not limited to whether the work should be performed but also to how it should be integrated into the existing workflow. Options may include reallocating crews from non-critical activities, resequencing work to avoid bottlenecks, or increasing capacity at constrained resources such as inspection or scaffolding. Without a system-level model, these decisions are often made in isolation, creating unintended downstream impacts that degrade overall performance.

 

Uncertainty also amplifies the importance of bottleneck management during execution. Emergent work tends to compete for the same constrained resources that govern planned work, such as inspection hold points or equipment isolation steps. If these bottlenecks are not actively managed, small-scope growth can produce disproportionate schedule impacts due to queue expansion. Real-time visibility into queue lengths and resource utilization allows teams to intervene early, either by redistributing work or temporarily increasing capacity, to maintain flow through the system.

 

In this context, contingency shifts from a passive allowance to an actively managed capability. A production model enables continuous re-forecasting of completion timelines as conditions evolve, allowing leadership to make informed trade-offs between scope, cost, and schedule. This transforms execution from a reactive process into a controlled system, where outcomes are driven by informed decisions rather than absorbed inefficiencies.

 

Ultimately, managing uncertainty in a turnaround is not about eliminating the unknown; it is about designing a system that can adapt to it. By treating execution as a dynamic production system and continuously aligning decisions with current system constraints, organizations can maintain schedule integrity, improve resource utilization, and ensure a safe and predictable return to service even in the presence of significant variability.


CONCLUSION


Production System Optimization transforms turnarounds from schedule-driven plans into flow-driven systems. By explicitly modeling variability, queueing effects, and resource constraints, organizations shift from reactive execution to controlled, data-driven delivery. This approach exposes true system bottlenecks, aligns work with available capacity, and enables more effective decision-making during execution.

Operators who apply these methods will consistently reduce schedule overruns, improve labor productivity, and accelerate safe return to service. More importantly, they gain the ability to manage uncertainty in real time, integrating emergent work without losing control of the overall system.

 

In an environment where every hour of downtime carries disproportionate economic impact, predictability is no longer a measure of operational excellence; it is a strategic competitive advantage.

references

 

 
 
bottom of page