C syntax
This article has multiple issues. Please help improve it or discuss these issues on the talk page. (Learn how and when to remove these template messages)
(Learn how and when to remove this template message)
|

C syntax is the form that text must have in order to be C programming language code. The language syntax rules are designed to allow for code that is terse, has a close relationship with the resulting object code, and yet provides relatively high-level data abstraction. C was the first widely successful high-level language for portable operating-system development.[1] C syntax makes use of the maximal munch principle. As a free-form language, C code can be formatted different ways without affecting its syntactic nature. C syntax influenced the syntax of succeeding languages, including C++, Java, and C#.
High level structure
C code consists of preprocessor directives, and core-language types, variables and functions, organized as one or more source files. Building the code typically involves preprocessing and then compiling each source file into an object file. Then, the object files are linked to create an executable image.
Variables and functions can be declared separately from their definition. A declaration identifies the name of a user-defined element and some if not all of the information about how the element can be used at run-time. A definition is a complete description of an element that includes the declaration aspect as well as additional information that completes the element. For example, a function declaration indicates the name and optionally the type and number of arguments that it accepts. A function definition includes the same information (argument information is not optional), plus code that implements the function logic.
Entry point

For a hosted environment, a program starts at an entry point function named main. The function is passed two arguments although an implementation of the function can ignore them. The function must be declared per one of the following prototypes (parameter names shown are typical but can be anything):
int main();
int main(void);
int main(int argc, char* argv[]);
int main(int argc, char** argv);
int main(void);
int main(int argc, char* argv[], char* envp[]);
int main(int argc, char** argv, char** envp);
int main(int argc, char* argv[], char** envp);
int main(int argc, char** argv, char* envp[]);
The first two definitions are equivalent, meaning that the function does not use the two arguments. The second two are also equivalent, allowing the function to access the two arguments which are needed to read command line arguments. The last four are also equivalent, allowing access to environment variables as well as command line arguments.
The return value, typed as int, serves as a status indicator to the host environment. Defined in <stdlib.h>, the standard library provides macros for standard status values: EXIT_SUCCESS and EXIT_FAILURE. Regardless, a program can indicate status using any values. For example, the kill command returns the numerical value of the signal plus 128.
A minimal program consists of an empty main function, like:
int main() {}
Unlike other functions, the language[lower-alpha 1] requires that a program act as if it returns 0 even if it does not end with a return statement.[2]
In a free-standing (non-hosted) environment, such as a system without an operating system, the standard allows for different startup handling. It need not require a main function.
Command-line arguments
Arguments included in the command line to start a program are passed to a program as two values – the number of arguments (customarily named argc) and an array of null-terminated strings (customarily named argv) with the program name as the first item.
The following code prints the value of the command-line parameters.
#include <stdio.h>
int main(int argc, char* argv[]) {
printf("argc\t= %d\n", argc);
for (int i = 0; i < argc; ++i) {
printf("argv[%i]\t= %s\n", i, argv[i]);
}
}
$ ./a.out abc def
argc = 3
argv[0] = ./a.out
argv[1] = abc
argv[2] = def
Reserved keywords
The following words are reserved – not allowed as identifiers, of which there are 43.
alignasalignofautoboolbreakcasecharconstconstexprcontinuedefaultdodoubleelseenumexternfloatforgotoifinlineintlongregisterrestrictreturnshortsignedsizeofstaticstatic_assertstructswitchthread_localtypedeftypeoftypeof_unqualunionunsignedvoidvolatilewhile
The following keywords are often substituted for a macro or an appropriate keyword from the above list, of which there are 14. Some of the following keywords are deprecated since C23.
_Alignas(deprecated)_Alignof(deprecated)_Atomic_BitInt_Bool(deprecated)_Complex_Countof_Decimal32_Decimal64_Decimal128_Generic_Noreturn(deprecated)_Static_assert(deprecated)_Thread_local(deprecated)
The keyword _Imaginary was removed in C2Y.
The following words refer to literal values used by the language, of which there are 3.
nullptrtruefalse
Implementations may reserve other keywords, although implementations typically provide non-standard keywords that begin with one or two underscores. The following keywords are classified as extensions and conditionally-supported, of which there are 2.
Preprocessor directives
The following are directives to the preprocessor, of which there are 19.
#if#elif#else#endif#ifdef#ifndef#elifdef#elifndef#define#undef#include#embed#line#error#warning#pragma#__has_include#__has_embed#__has_c_attribute
The _Pragma operator provides an alternative syntax for the functionality provided by #pragma.
Comments
A comment – informative text to a programmer that is ignored by a language translator – can be included in code either as the line or block comment syntax. A line comment starts with // and ends at the end of the same line. A block comment starts with /* and ends with */ – spanning any number of lines (or just one).
In some situations, comment markers are ignored. The text of a string literal is exempt from being considered a comment start. And, comments cannot be nested. For example, /* in a line comment is not as treated as the start of a block comment. And, // in a block comment is not treated as the start of a line comment.
The line comment syntax, sometimes called C++ style originated in BCPL and became valid syntax in C99. It is not available in the original K&R version nor in ANSI C.
The following code demonstrates comments. Line 1 contains a line comment and lines 3-4 contain a block comment. Line 4 demonstrates that a block comment can be embedded in a line with code both before and after it.
int i; // line comment
/* block
comment */
int ii = /* always zero */ 0;
The following demonstrates a potential problem with the comment syntax. What is intended to be the divide operator / and then the dereference operator *, is evaluated as the start of a block comment.
x = *p/*q;
The following text is not valid C syntax since comments do not nest. It seems that lines 3-5 are a comment nested inside the comment block spanning lines 1-7. But, actually line 5 ends the comment started on line 1. This leaves line 6 to be interpreted as code which is clearly not valid C syntax.
/*
First of comment block
/*
First line of what is intended to be an inner block
*/
Compiler treats this line as code but it's not valid!
*/
Identifiers
The syntax supports user-defined identifiers. An identifier must start with a letter (A-Z, a-z) or an underscore (_), subsequent characters can be letters, numbers (0-9), or underscores and it must not be a reserved word. Identifiers are case sensitive, making foo, FOO, and Foo distinct.
Evaluation order
There can be multiple ways to evaluate an expression consistent with the mathematical notation. For example, (1+1)+(3+3) may be evaluated in the order (1+1)+(3+3), (2)+(3+3), (2)+(6), (8), or in the order (1+1)+(3+3), (1+1)+(6), (2)+(6), (8).
To reduce run-time issues with evaluation order yet afford some optimizations, the standard states that expressions may be evaluated in any order between sequence points which are defined as any of the following:
- Statement end
- A sequencing operator: a comma; commas that delimit function arguments are not sequence points
- A short-circuit operators: logical and (
&&, which can be read and then) and logical or (||, which can be read or else) - A ternary operator (
?:): This operator evaluates its first sub-expression first, and then its second or third (never both of them) based on the value of the first - Entry to and exit from a function call (but not between evaluations of the arguments)
Expressions before a sequence point are always evaluated before those after a sequence point. In the case of short-circuit evaluation, the second expression may not be evaluated depending on the result of the first expression. For example, in the expression (a() || b()), if the first argument evaluates to nonzero (true), the result of the entire expression cannot be anything else than true, so b() is not evaluated. Similarly, in the expression (a() && b()), if the first argument evaluates to zero (false), the result of the entire expression cannot be anything else than false, so b() is not evaluated.
The arguments to a function call may be evaluated in any order, as long as they are all evaluated by the time the function is entered. The following expression, for example, has undefined behavior:
printf("%s %s\n", argv[i = 0], argv[++i]);
Code inclusion
Headers
Code is included from other files by using the #include directive, which textually copies the contents of the file in-place. Traditionally, C code is divided between a header file (with extension .h) and a source file (with extension .c). The header contains the declarations of symbols, while the source file contains the full implementation. This separation is enforced to prevent the compiler from recompiling the same source code repeatedly every time a library is included, causing bloat to compile times. Headers and separate translation units have no sense of namespaces, meaning all symbol names are global and will clash if the same name for multiple objects exist in the same translation unit.
To prevent a header from being included into a file more than once, #include guards or #pragma once can be used.
C headers can be compiled into precompiled headers which are faster to process by the compiler. Typically, libraries that infrequently change (such as C standard library headers) could be precompiled for faster compilation in a project.
To distinguish between searching in an include directory and searching a relative path, use angle brackets for include directories and quotation marks for relative paths.
#pragma once
// standard library includes
#include <stdio.h>
#include <stdlib.h>
// local
#include "Personalities.h"
#include "Math/SpecialFunctions.h"
// external libraries
#include <sqlite3.h>
Clang C modules
Clang offers a non-standard feature, called modules, which are similar to C++ modules but different semantically. Both are used to increase compile time by compiling a translation unit once.
Embed
The #embed directive can be used to embed binary content into a file, even if it is not valid C code.
constexpr char ICON_DISPLAY_DATA[] = {
#embed "art.png"
};
// specify any type which can be initialized form integer constant expressions will do
constexpr char RESET_BLOB[] = {
#embed "data.bin"
};
// attributes work just as well
alignas(8) constexpr char ALIGNED_DATA_STRING[] = {
#embed "attributes.xml"
};
int main() {
return
#embed </dev/urandom> limit(1)
;
}
Type system
Primitive types
The language supports primitive numeric types for integer and real values which typically map directly to the instruction set architecture of a central processing unit (CPU). Integer data types store values in a subset of integers, and real data types store values in a subset of real numbers in floating-point. A complex data type stores two real values.
Integer types have signed and unsigned variants. If neither is specified, signed is assumed, in most circumstances. However, for historic reasons, char is a type distinct from both signed char and unsigned char. It may be signed or unsigned, depending on the compiler and the character set (the standard requires that members of the basic character set have positive values). Also, bit field types specified as int may be signed or unsigned, depending on the compiler.
Integer types
The integer types come in different fixed sizes, capable of representing various ranges of numbers. The type char occupies exactly one byte (the smallest addressable storage unit), which is typically 8 bits wide. (Although char can represent any "basic" character, a wider type may be required for international character sets.) Most integer types have both signed and unsigned varieties, designated by the signed and unsigned keywords. Signed integer types always use the two's complement representation, since C23[3] (and in practice before; in versions before C23 the representation might alternatively have been ones' complement, or sign-and-magnitude, but in practice that has not been the case for decades on modern hardware). In many cases, there are multiple equivalent ways to designate the type; for example, signed short int and short are synonymous.
The representation of some types may include unused "padding" bits, which occupy storage but are not included in the width. The following table lists the integer types using the shortest possible name and indicating the minimum width in bits.
| Name | Minimum width (bits) |
|---|---|
| style="text-align: center" | 1 | |
char
|
8 |
signed char
|
8 |
unsigned char
|
8 |
short
|
16 |
unsigned short
|
16 |
int
|
16 |
unsigned int
|
16 |
long
|
32 |
unsigned long
|
32 |
long long[note 1]
|
64 |
unsigned long long[note 1]
|
64 |
The char type is distinct from both signed char and unsigned char, but is guaranteed to have the same representation as one of them. The _Bool and long long types are standardized since 1999, and may not be supported by older compilers. Type _Bool is usually accessed via the typedef name bool defined by the standard header <stdbool.h>, however since C23 the _Bool type has been renamed bool, and <stdbool.h> has been deprecated.
In general, the widths and representation scheme implemented for any given platform are chosen based on the machine architecture, with some consideration given to the ease of importing source code developed for other platforms. The width of the int type varies especially between translators; it often corresponds to the most "natural" word size for a platform. The standard header <limits.h> defines macros for the minimum and maximum representable values of the standard integer types as implemented on any specific platform.
In addition to the standard integer types, there may be other "extended" integer types, which can be used for typedefs in standard headers. For more precise specification of width, programmers can and should use typedefs from the standard header <stdint.h>.
Integer constants may be specified in source code in several ways. Numeric values can be specified as decimal (example: 1022), octal with zero (0) as a prefix (01776), or hexadecimal with 0x (zero x) as a prefix (0x3FE). A character in single quotes (example: 'R'), called a "character constant," represents the value of that character in the execution character set, with type int. Except for character constants, the type of an integer constant is determined by the width required to represent the specified value, but is always at least as wide as int. This can be overridden by appending an explicit length and/or signedness modifier; for example, 12lu has type unsigned long. There are no negative integer constants, but the same effect can often be obtained by using a unary negation operator "-".
Enumerated type
The enumerated type, specified with the enum keyword, and often just called an "enum" (usually pronounced /ˈiːnʌm/ EE-num or /ˈiːnuːm/ EE-noom), is a type designed to represent values across a series of named constants. Each of the enumerated constants has type int. Each enum type itself is compatible with char or a signed or unsigned integer type, but each implementation defines its own rules for choosing a type.
Some compilers warn if an object with enumerated type is assigned a value that is not one of its constants. However, such an object can be assigned any values in the range of their compatible type, and enum constants can be used anywhere an integer is expected. For this reason, enum values are often used in place of preprocessor #define directives to create named constants. Such constants are generally safer to use than macros, since they reside within a specific identifier namespace.
An enumerated type is declared with the enum specifier and an optional name (or tag) for the enum, followed by a list of one or more constants contained within curly braces and separated by commas, and an optional list of variable names. Subsequent references to a specific enumerated type use the enum keyword and the name of the enum. By default, the first constant in an enumeration is assigned the value zero, and each subsequent value is incremented by one over the previous constant. Specific values may also be assigned to constants in the declaration, and any subsequent constants without specific values will be given incremented values from that point onward.
For example, consider the following declaration:
enum Color {
RED,
GREEN,
BLUE = 5,
YELLOW
} paint_color;This declares the enum Color type; the int constants RED (whose value is 0), GREEN (whose value is one greater than RED, 1), BLUE (whose value is the given value, 5), and YELLOW (whose value is one greater than BLUE, 6); and the enum Color variable paint_color. The constants may be used outside of the context of the enum (where any integer value is allowed), and values other than the constants may be assigned to paint_color, or any other variable of type enum Color.
Unlike C++, C enums are not scoped, as C has no concept of namespacing. In C, enums can be implicitly converted to numeric types, which is type-unsafe.
typedef enum Color {
RED,
ORANGE,
YELLOW,
GREEN,
BLUE,
INDIGO,
VIOLET
} Color;
Color c = RED; // in C
Color d = Color::RED; // in C++, but not in CSince C23, it is possible to manually specify the underlying type for the enum, like in C++.
enum CardSuit: char {
HEARTS,
CLUBS,
SPADES,
DIAMONDS
};Floating-point types
A floating-point form is used to represent numbers with a fractional component. They do not, however, represent most rational numbers exactly; they are instead a close approximation. There are three standard types of real values, denoted by their specifiers (and since C23 three more decimal types): single precision (float), double precision (double), and double extended precision (long double). Each of these may represent values in a different form, often one of the IEEE floating-point formats.
| Type specifiers | Precision (decimal digits) | Exponent range | ||
|---|---|---|---|---|
| Minimum | IEEE 754 | Minimum | IEEE 754 | |
float
|
6 | 7.2 (24 bits) | ±37 | ±38 (8 bits) |
double
|
10 | 15.9 (53 bits) | ±37 | ±307 (11 bits) |
long double
|
10 | 34.0 (113 bits) | ±37 | ±4931 (15 bits) |
Floating-point constants may be written in decimal notation, e.g. 1.23. Decimal scientific notation may be used by adding e or E followed by a decimal exponent, also known as E notation, e.g. 1.23e2 (which has the value 1.23 × 102 = 123.0). Either a decimal point or an exponent is required (otherwise, the number is parsed as an integer constant). Hexadecimal floating-point constants follow similar rules, except that they must be prefixed by 0x and use p or P to specify a binary exponent, e.g. 0xAp-2 (which has the value 2.5, since Ah × 2−2 = 10 × 2−2 = 10 ÷ 4). Both decimal and hexadecimal floating-point constants may be suffixed by f or F to indicate a constant of type float, by l (letter l) or L to indicate type long double, or left unsuffixed for a double constant.
The standard header file <float.h> defines the minimum and maximum values of the implementation's floating-point types float, double, and long double. It also defines other limits that are relevant to the processing of floating-point numbers.
C23 introduces three additional decimal (as opposed to binary) real floating-point types: _Decimal32, _Decimal64, and _Decimal128.
- NOTE C does not specify a radix for
float,double, andlong double. An implementation can choose the representation offloat,double, andlong doubleto be the same as the decimal floating types.[4]
Despite that, the radix has historically been binary (base 2), meaning numbers like 1/2 or 1/4 are exact, but not 1/10, 1/100 or 1/3. With decimal floating point all the same numbers are exact plus numbers like 1/10 and 1/100, but still not e.g. 1/3. No known implementation does opt into the decimal radix for the previously known to be binary types. Since most computers do not even have the hardware for the decimal types, and those few that do (e.g. IBM mainframes since IBM System z10), can use the explicitly decimal types.
Storage class
The following table describes the specifiers that define various storage attributes including duration – static (default for global), automatic (default for local), or dynamic (allocated).[citation needed]
| Specifier | Lifetime | Scope | Default initializer |
|---|---|---|---|
auto
|
Block (stack) | Block | Uninitialized |
register
|
Block (stack or CPU register) | Block | Uninitialized |
static
|
Program | Block or compilation unit | Zero |
extern
|
Program | Global (entire program) | Zero |
thread_local
|
Thread | ||
| (none)1 | Dynamic (heap) | Uninitialized (initialized to 0 if using calloc())
|
- 1 Allocated and deallocated using the
malloc()andfree()library functions.
Variables declared within a block by default have automatic storage, as do those explicitly declared with the auto[note 2] or register storage class specifiers. The auto and register specifiers may only be used within functions and function argument declarations;[citation needed] as such, the auto specifier is always redundant. Objects declared outside of all blocks and those explicitly declared with the static storage class specifier have static storage duration. Static Variables are initialized to zero by default by the compiler.[citation needed]
Objects with automatic storage are local to the block in which they were declared and are discarded when the block is exited. Additionally, objects declared with the register storage class may be given higher priority by the compiler for access to registers, although the compiler may choose not to actually store any of them in a register. Objects with this storage class may not be used with the address-of (&) unary operator. Objects with static storage persist for the program's entire duration. In this way, the same object can be accessed by a function across multiple calls. Objects with allocated storage duration are created and destroyed explicitly with malloc, free, and related functions.
The extern storage class specifier indicates that the storage for an object has been defined elsewhere. When used inside a block, it indicates that the storage has been defined by a declaration outside of that block. When used outside of all blocks, it indicates that the storage has been defined outside of the compilation unit. The extern storage class specifier is redundant when used on a function declaration. It indicates that the declared function has been defined outside of the compilation unit.
The thread_local (_Thread_local before C23,[citation needed] and in earlier versions of C if the header <threads.h> is included) storage class specifier, introduced in C11, is used to declare a thread-local variable. It can be combined with static or extern to determine linkage.
Note that storage specifiers apply only to functions and objects; other things such as type and enum declarations are private to the compilation unit in which they appear.[citation needed] Types, on the other hand, have qualifiers (see below).
Since C23, C can use auto to declare a type-inferred variable.
Type qualifiers
Types can be qualified to indicate special properties of their data. The type qualifier const indicates that a value does not change once it has been initialized. Attempting to modify a const qualified value yields undefined behavior, so some compilers store them in rodata or (for embedded systems) in read-only memory (ROM). Similarly, constexpr can be thought of as a "stronger" form of const, where the value must be known at compile time (making it a type-safe replacement for macro constants). A constexpr function, similarly, must also be able to be evaluated at compile time. The type qualifier volatile indicates to an optimizing compiler that it may not remove apparently redundant reads or writes, as the value may change even if it was not modified by any expression or statement, or multiple writes may be necessary, such as for memory-mapped I/O.
Incomplete types
An incomplete type is a structure or union type whose members have not yet been specified, an array type whose dimension has not yet been specified, or the void type (the void type cannot be completed). Such a type may not be instantiated (its size is not known), nor may its members be accessed (they, too, are unknown); however, the derived pointer type may be used (but not dereferenced).
They are often used with pointers, either as forward or external declarations. For instance, code could declare an incomplete type like this:
struct Integer* pt;This declares pt as a pointer to struct Integer (as well as the incomplete struct type). As all pointers have the same size (regardless of what they point to), code can use pt as a pointer although it cannot access the fields of struct Integer.
An incomplete type can be completed later in the same scope by redeclaring it. For example:
struct Integer {
int num;
};Incomplete types are used to implement recursive structures; the body of the type declaration may be deferred to later in the translation unit:
typedef struct Bert Bert;
typedef struct Wilma Wilma;
struct Bert {
Wilma* wilma;
};
struct Wilma {
Bert* bert;
};Incomplete types are also used for data hiding. The incomplete type is defined in a header file, and the full definition is hidden in a single body file.
Pointers
In a variable declaration, the asterisk (*) can be considered to mean "pointer-to". For example, int x declares a variable of type int, and int* px declares a variable px that is a pointer to integer. Some[who?] contend that based on the language definition, the * is more closely related to the variable than the type and therefore format the code as int *px or even int * px. This is because the asterisk(*) is part of the declarator[5] and not the type specifier; another consequence of this is that int* a,b; will declare two variables such that a is a pointer to an integer and b is an integer itself.
A pointer value associates two pieces of information: a memory address and a data type.
Referencing
When a non-static pointer is declared, it has an unspecified value. Dereferencing it without first assigning it, results in undefined behavior.
The & operator specifies the address of the data object after it. In the following example, ptr is assigned the address of a:
int a = 0;
int* ptr = &a;Dereferencing
An asterisk before a variable name (when not in a declaration or a mathematical expression) dereferences a pointer to allow access to the value it points to. In the following example, the integer variable b is set to the value of integer variable a, which is 10:
int a = 10;
int* p;
p = &a;
int b = *p;Arrays
Array definition
Arrays store consecutive elements of the same type. The following code declares an array of 100 elements, named a, of type int.
int a[100];If declared outside of a function (globally), the size must be a constant value. If declared in a function, the array size may be a non-constant expression.
The number of elements is available as sizeof(a) / sizeof(int), but if the value is passed to another function, the number of elements is not available via the formal parameter variable.
Accessing elements
The primary facility for accessing array elements is the array subscript operator. For example, a[i] accesses the element at index i of array a. Array indexing begins at 0, making the last array index equal to the number of elements minus 1. As the standard does not provide for array indexing bounds checking, specifying an index that is out of range, results in undefined behavior.
Due to arrays and pointers being interchangeable, the address of each elements can be expressed in pointer arithmetic. The following table illustrates both methods for the existing the same array:
| Element | First | Second | Third | nth |
|---|---|---|---|---|
| Array subscript | a[0]
|
a[1]
|
a[2]
|
a[n - 1]
|
| Dereferenced pointer | *a
|
*(a + 1)
|
*(a + 2)
|
*(a + n - 1)
|
Since the expression a[i] is semantically equivalent to *(a + i), which in turn is equivalent to *(i + a), the expression can also be written as i[a], although this form is rarely used.
Variable-length arrays
C99 standardized the variable-length array (VLA) in block scope that produced an array sized by runtime information (not a constant value) but with fixed size until the end of the block.[2] As of C11, this feature is no longer required to be implemented by the compiler.
int n = 20; // n can be any arbitrary value
int a[n];
a[3] = 10;Multidimensional arrays
The language supports arrays of multiple dimensions – stored in row-major order which is essentially a one-dimensional array with elements that are arrays. Given that ROWS and COLUMNS are constants, the following declares a two-dimensional array of length ROWS, each element of which is an array of COLUMNS integers.
int array2d[ROWS][COLUMNS];The following is an example of accessing an integer element:
array2d[4][3]Reading from left to right, this accesses the 5th row, and the 4th element in that row. The expression array2d[4] is an array, which is then subscripting with [3] to access the fourth integer.
| Element | First | Second row, second column | ith row, jth column |
|---|---|---|---|
| Array subscript | array[0][0]
|
array[1][1]
|
array[i - 1][j - 1]
|
| Dereferenced pointer | *(*(array + 0) + 0)
|
*(*(array + 1) + 1)
|
*(*(array + i - 1) + j - 1)
|
Higher-dimensional arrays can be declared in a similar manner.
A multidimensional array should not be confused with an array of pointers to arrays (also known as an Iliffe vector or sometimes an array of arrays). The former is always rectangular (all subarrays must be the same size), and occupies a contiguous region of memory. The latter is a one-dimensional array of pointers, each of which may point to the first element of a subarray in a different place in memory, and the sub-arrays do not have to be the same size.
Text
Although the language provides types for textual character data, neither the language nor the standard library defines a string type, but the null terminated string is commonly used. A string value is a contiguous series of characters with the end denoted by a zero value. The standard library contains many string handling functions for null-terminated strings, but string manipulation can and often is handled via custom code.
String literal
A string literal is code text surrounded by double quotes, such as "Hello world!". A literal compiles to an array of the specified char values with a terminating null terminating character to mark the end of the string.
The language supports string literal concatenation – adjacent string literals are treated as joined at compile time. This allows long strings to be split over multiple lines, and also allows string literals from preprocessor macros to be appended to strings at compile time. For example, the source code:
printf(__FILE__ ": %d: Hello "
"world\n");becomes the following after the preprocessor expands __FILE__:
printf("helloworld.c" ": %d: Hello "
"world\n");which is equivalent to:
printf("helloworld.c: %d: Hello world\n");Character constants
The character literal, called character constant, is single-quoted, e.g. 'A', and has type int. To illustrate the difference between a string literal and a character constant, consider that "A" is two characters, 'A' and '\0', whereas 'A' represents a single character (65 in ASCII).
A character constant cannot be empty (i.e. '' is invalid syntax). Multi-character constants (e.g. 'xy') are valid, although rarely useful — they let one store several characters in an integer (e.g. 4 ASCII characters can fit in a 32-bit integer, 8 in a 64-bit one). Since the order in which the characters are packed into an int is not specified (left to the implementation to define), portable use of multi-character constants is difficult.
Nevertheless, in situations limited to a specific platform and the compiler implementation, multicharacter constants do find their use in specifying signatures. One common use case is the OSType, where the combination of Classic Mac OS compilers and its inherent big-endianness means that bytes in the integer appear in the exact order of characters defined in the literal. The definition by popular "implementations" are in fact consistent: in GCC, Clang, and Visual C++, '1234' yields 0x31323334 under ASCII.[7][8]
Like string literals, character constants can also be modified by prefixes, for example L'A' has type wchar_t and represents the character value of "A" in the wide character encoding.
Backslash escapes
Control characters cannot be included in a string or character literal directly. Instead they can be encoded via an escape sequence starting with a backslash (\). For example, the backslashes in "This string contains \"double quotes\"." indicate that the inner pair of quotes are intended as an actual part of the string, rather than the default reading as a delimiter (endpoint) of the string.
Escape sequences include:
| Sequence | Meaning |
|---|---|
\\ |
Literal backslash |
\" |
Double quote |
\' |
Single quote |
\n |
Newline (line feed) |
\r |
Carriage return |
\b |
Backspace |
\t |
Horizontal tab |
\f |
Form feed |
\a |
Alert (bell) |
\v |
Vertical tab |
\? |
Question mark (used to escape trigraphs, obsolete feature dropped in C23) |
\OOO |
Character with octal value OOO (where OOO is 1-3 octal digits, '0'-'7') |
\xhh |
Character with hexadecimal value hh (where hh is 1 or more hex digits, '0'-'9','A'-'F','a'-'f') |
\uhhhh |
Unicode code point below 10000 hexadecimal (added in C99) |
\Uhhhhhhhh |
Unicode code point where hhhhhhhh is eight hexadecimal digits (added in C99) |
The use of other backslash escapes is not defined by the standard, although compilers often provide additional escape codes as language extensions. For example, the escape sequence \e for the escape character with ASCII hex value 1B which was not added to the standard due to lacking representation in other character sets (such as EBCDIC). It is available in GCC, clang and tcc.
Note that the standard library function printf() uses %% to represent the literal % character.
Wide character strings
Since type char is 1 byte wide, a single char value typically can represent at most 255 distinct character codes, not nearly enough for all the different characters in use worldwide. To provide better support for international characters, the first standard (C89) introduced wide characters (encoded in type wchar_t) and wide character strings, which are written as L"Hello world!"
Wide characters are most commonly either 2 bytes (using a 2-byte encoding such as UTF-16) or 4 bytes (usually UTF-32), but Standard C does not specify the width for wchar_t, leaving the choice to the implementor. Microsoft Windows generally uses UTF-16, thus the above string would be 26 bytes long for a Microsoft compiler; the Unix world prefers UTF-32, thus compilers such as GCC would generate a 52-byte string. A 2-byte wide wchar_t suffers the same limitation as char, in that certain characters (those outside the BMP) cannot be represented in a single wchar_t, and must be represented using surrogate pairs.
The original standard specified only minimal functions for operating with wide character strings; in 1995 the standard was modified to include much more extensive support, comparable to that for char strings. The relevant functions are mostly named after their char equivalents, with the addition of a "w" or the replacement of "str" with "wcs"; they are specified in <wchar.h>, with <wctype.h> containing wide-character classification and mapping functions.
The now generally recommended method[note 3] of supporting international characters is through UTF-8, which is stored in char arrays, and can be written directly in the source code if using a UTF-8 editor, because UTF-8 is a direct ASCII extension.
Variable width strings
A common alternative to wchar_t is to use a variable-width encoding, whereby a logical character may extend over multiple positions of the string. Variable-width strings may be encoded into literals verbatim, at the risk of confusing the compiler, or using numerical backslash escapes (e.g. "\xc3\xa9" for "é" in UTF-8). The UTF-8 encoding was specifically designed (under Plan 9) for compatibility with the standard library string functions; supporting features of the encoding include a lack of embedded nulls, a lack of valid interpretations for subsequences, and trivial resynchronisation. Encodings lacking these features are likely to prove incompatible with the standard library functions; encoding-aware string functions are often used in such cases.
Structure
A structure or struct is a container consisting of a sequence of named members of heterogeneous types, similar to a record in other languages. The first field starts at the address of the structure and the members are stored in consecutive locations in memory, but the compiler can insert padding between or after members for efficiency or as padding required for proper alignment by the target architecture. The size of a structure includes padding.
A structure is declared with the struct keyword followed by an optional identifier name, which is used to identify the form of the structure. The body follows with field declarations that each consist of a type name, a field name and terminated with a semi-colon.
The following declares a structure named MyStruct that contains three members. It also declares an instance named tee:
struct MyStruct {
int x;
float y;
char* z;
} tee;Structure members cannot have an incomplete or function type. Thus members cannot be an instance of the structure being declared (because it is incomplete at that point) but a field can be a pointer to the type being declared.
Once declared, a variable can be declared of the structure type.
The following declares a new instance of the structure MyStruct named r:
struct MyStruct r;Although some prefer to declare a struct variable using the struct keyword, some use typedef to alias the struct type into the main type namespace. The following declares a type as Integer which can then be used like Integer n.
typedef struct {
int i;
} Integer;Accessing members
A member is accessed using dot notation. For example, given the declaration of tee from above, the member y can be accessed as tee.y.
A structure is commonly accessed via a pointer. Consider struct MyStruct* ptee = &tee that defines a pointer to tee, named ptee. Member y of tee can be accessed by dereferencing ptee and using the result as the left operand as (*ptee).y.[lower-alpha 2] Because this operation is common, the language provides an abbreviated syntax for accessing a member directly from a pointer, -> (for example, ptee->y).
Assignment
Assigning a value to a member is like assigning a value to a variable. The only difference is that the lvalue (left side value) of the assignment is the name of the member according to the above syntax.
A structure can also be assigned as a whole to another structure of the same type, passed by copy as a function argument or return value. For example, tee.x = 74 assigns the value 74 to the member named x in the structure tee, And, ptee->x = 74 does the same for ptee.
Other operations
The operations supported for a structure are: initialize, copy, get address and access a field. Of note, the language does not support comparing the value of two structures other than via custom code to compare each field.
Bit fields
The language provides a special type of member known as a bit field, which is an integer with a specified size in bits. A bit field is declared as a member of type (signed/unsigned) int, bool, or _BitInt(N).[note 4] plus a suffix after the member name consisting of a colon and a number of bits. The total number of bits in a single bit field must not exceed the total number of bits of its base type. Contrary to the usual C syntax rules, it is implementation-defined whether a bit field is signed or unsigned if not explicitly specified. Therefore, best practice is to specify signed or unsigned.
Unnamed fields indicate padding and consist of just a colon followed by a number of bits. Specifying a width of zero for an unnamed field is used to force alignment to a new word.[9] Since all members of a union occupy the same memory, unnamed bit-fields of width zero do nothing in unions, however unnamed bit-fields of non zero width can change the size of the union since they have to fit in it.
Bit fields are limited compared to normal fields in that the address-of (&) and sizeof operators are not supported.
The following declares a structure type named FlagStatus and an instance of it named g. The first field, flag, is a single bit flag, which can only be 1 or 0. The second field, num, is a signed 4-bit field; range -7...7 or -8...7. The last field adds 3 bits of padding to round out the structure to 8 bits.
struct FlagStatus {
unsigned int flag : 1;
signed int num : 4;
signed int : 3;
} g;Namespaces
C itself has no native support for namespaces unlike C++ and Java. This makes C symbol names prone to name clashes. However, it is possible to use anonymous structs to emulate namespaces.
Math.h:
#pragma once
const struct {
double PI;
double (*sin)(double);
} Math;Math.c:
#include <math.h>
static double _sin(double arg) {
return sin(arg);
}
const struct {
double PI;
double (*sin)(double);
} Math = { M_PI, _sin };Main.c:
#include <stdio.h>
#include "Math.h"
int main() {
printf("sin(0) = %d\n", Math.sin(0));
printf("pi is %f\n", Math.PI);
}Union
For the most part, a union is like a structure except that fields overlap in memory to allow storing values of different type although not at the same time. The union is like the variant record of other languages. Each field refers to the same location in memory. The size of a union is equal to the size of its largest component type plus any padding.
A union is declared with the union keyword. The following declares a union named MyUnion and an instance of it named n:
union MyUnion {
int x;
float y;
char* z;
} n;Initialization
Scalar
Initializing a variable along with declaring it involves appending an equals sign and then a construct that is compatible with the data type. The following initializes an int:
int x = 12;Because of the language's grammar, a scalar initializer may be enclosed in any number of curly brace pairs. Most compilers issue a warning if there is more than one such pair. The following are legal although arguably unusual:
int y = { 23 };
int z = { { 34 } };Initializer list
Structures, unions and arrays can be initialized after a declaration via an initializer list.
Since unmatched elements are set to 0, an empty list sets all elements to 0. For example, the following sets all elements of array a and all fields of s to 0:
int a[10] = {};
struct MyStruct s = {};If an array is declared without an explicit size, the array is an incomplete type. The number of initializers determines the size of the array and completes the type. For example:
int x[] = { 0, 1, 2 };By default, the items of an initializer list correspond with the elements in the order they are defined. Including too many values yields an error. The following statement initializes an instance of the structure MyStruct named pi:
struct MyStruct {
int x;
float y;
char* z;
};
struct MyStruct pi = { 3, 3.1415, "Pi" };Designated initializers
Designated initializers allow members to be initialized by name, in any order, and without explicitly providing preceding values. The following initialization is functionally equivalent to the previous:
struct MyStruct pi = { .z = "Pi", .x = 3, .y = 3.1415 };Using a designator in an initializer moves the initialization "cursor". In the example below, if MAX is greater than 10, there will be some zero-valued elements in the middle of a; if it is less than 10, some of the values provided by the first five initializers will be overridden by the second five. If MAX is less than 5, there will be a compilation error:
int a[MAX] = { 1, 3, 5, 7, 9, [MAX - 5] = 8, 6, 4, 2, 0 };In C89, a union was initialized with a single value applied to its first member. That is, the union MyUnion defined above could only have its x member initialized:
union MyUnion value = { 3 };Using a designated initializer, the member to be initialized does not have to be the first member:
union MyUnion value = { .y = 3.1415 };Compound designators can be used to provide explicit initialization when unadorned initializer lists might be misunderstood. In the example below, w is declared as an array of structures, each structure consisting of a member a (an array of 3 int) and a member b (an int). The initializer sets the size of w to 2 and sets the values of the first element of each a:
struct { int a[3], b; } w[] = { [0].a = {1}, [1].a[0] = 2 };This is equivalent to:
struct { int a[3], b; } w[] =
{
{ { 1, 0, 0 }, 0 },
{ { 2, 0, 0 }, 0 }
};Compound literals
It is possible to borrow the initialization methodology to generate compound structure and array literals:
// pointer created from array literal.
int* ptr = (int[]){ 10, 20, 30, 40 };
// pointer to array.
float (*foo)[3] = &(float[]){ 0.5f, 1.f, -0.5f };
struct MyStruct pi = (struct MyStruct){ 3, 3.1415, "Pi" };Compound literals are often combined with designated initializers to make the declaration more readable:[2]
pi = (struct MyStruct){ .z = "Pi", .x = 3, .y = 3.1415 };Function pointers
A pointer to a function can be declared like:
type-name (*function-name)(parameter-list);
The following program code demonstrates the use of a function pointer for selecting between addition and subtraction. Line 12 defines a function pointer variable named operation that supports the same interface as both the add and subtract functions. Based on the conditional (argc), operation is assigned to the address of either add or subtract. On line 14, the function that operation points to is called.
nodiscard
constexpr int add(int x, int y) {
return x + y;
}
nodiscard
constexpr int subtract(int x, int y) {
return x - y;
}
int main(int argc, char* argv[]) {
int (*operation)(int x, int y);
operation = argc ? add : subtract;
int result = operation(1, 1);
return result;
}Operators

This is a list of operators in the C and C++ programming languages. All the operators (except typeof) listed exist in C++; the column "Included in C", states whether an operator is also present in C. Note that C does not support operator overloading.
When not overloaded, for the operators &&, ||, and , (the comma operator), there is a sequence point after the evaluation of the first operand.
C++ also contains the type conversion operators const_cast, static_cast, dynamic_cast, and reinterpret_cast. The formatting of these operators means that their precedence level is unimportant.
Most of the operators available in C and C++ are also available in other C-family languages such as C#, D, Java, Perl, and PHP with the same precedence, associativity, and semantics.
Table
For the purposes of these tables, a, b, and c represent valid values (literals, values from variables, or return value), object names, or lvalues, as appropriate. R, S and T stand for any type(s), and K for a class type or enumerated type.
Arithmetic operators
All arithmetic operators exist in C and C++ and can be overloaded in C++.
Comparison operators/relational operators
All comparison operators can be overloaded in C++. Since C++20, the inequality operator is automatically generated if operator== is defined and all four relational operators are automatically generated if operator<=> is defined.[10]
Logical operators
All logical operators exist in C and C++ and can be overloaded in C++, albeit the overloading of the logical AND and logical OR is discouraged, because as overloaded operators they behave as ordinary function calls, which means that both of their operands are evaluated, so they lose their well-used and expected short-circuit evaluation property.[11]
Bitwise operators
All bitwise operators exist in C and C++ and can be overloaded in C++.
Assignment operators
All assignment expressions exist in C and C++ and can be overloaded in C++.
For the given operators the semantic of the built-in combined assignment expression a ⊚= b is equivalent to a = a ⊚ b, except that a is evaluated only once.
Member and pointer operators
| Operator name | Syntax | Can overload in C++ | Included in C |
C++ prototype examples | ||
|---|---|---|---|---|---|---|
| As member of K | Outside class definitions | |||||
| Subscript | a[b]a<:b:>a??(b??)[lower-alpha 3][lower-alpha 4]
|
Yes | Yes | R& K::operator [](S b);R& K::operator [](S b, ...); // since C++23 |
N/A | |
| Indirection ("object pointed to by a") | *a |
Yes | Yes | R& K::operator *();
|
R& operator *(K a);
| |
| Address-of ("address of a") | &abitand a[lower-alpha 5][lower-alpha 6] |
Yes[lower-alpha 7] | Yes | R* K::operator &();
|
R* operator &(K a);
| |
| Structure dereference ("member b of object pointed to by a") | a->b |
Yes | Yes | R* K::operator ->();[lower-alpha 8] |
N/A | |
| Structure reference ("member b of object a") | a.b |
No | Yes | N/A | ||
| Member selected by pointer-to-member b of object pointed to by a[lower-alpha 9] | a->*b |
Yes | No | R& K::operator ->*(S b);
|
R& operator ->*(K a, S b);
| |
| Member of object a selected by pointer-to-member b | a.*b |
No | No | N/A | ||
Other operators
| Operator name | Syntax | Can overload in C++ | Included in C |
Prototype examples | ||
|---|---|---|---|---|---|---|
| As member of K | Outside class definitions | |||||
| Function call See Function object. |
a(a1, a2)
|
Yes | Yes | R K::operator ()(S a, T b, ...);
|
N/A | |
| Comma | a, b |
Yes | Yes | R K::operator ,(S b);
|
R operator ,(K a, S b);
| |
| Ternary conditional | a ? b : c |
No | Yes | N/A | ||
| Scope resolution | a::b |
No | No | N/A | ||
| User-defined literals[lower-alpha 10] since C++11 |
"a"_b |
Yes | No | N/A | R operator "" _b(T a)
| |
| Sizeof | sizeof a[lower-alpha 11]sizeof (R) |
No | Yes | N/A | ||
| Size of parameter pack since C++11 |
sizeof...(Args) |
No | No | N/A | ||
| Alignof since C++11 |
alignof(R) or _Alignof(R)[lower-alpha 12] |
No | Yes | N/A | ||
| Typeof since C23 |
typeof (a)typeof (R)typeof_unqual (a)typeof_unqual (R) |
N/A | Only[lower-alpha 13] | N/A | ||
| Decltype since C++11 |
decltype (a)decltype (R) |
No | No | N/A | ||
| Type identification | typeid(a)typeid(R) |
No | No | N/A | ||
| Conversion (C-style cast) | (R)a |
Yes | Yes | K::operator R();[12]
|
N/A | |
| Conversion | R(a)R{a}since C++11auto(a)since C++23auto{a}since C++23 |
No | No | Note: behaves like const_cast/static_cast/reinterpret_cast. In the last two cases the auto specifier is replaced with the type of the invented variable x declared with auto x(a); (which is never interpreted as a function declaration) or auto x{a};, respectively. [13]
| ||
| static cast conversion | static_cast<R>(a) |
Yes | No | K::operator R();explicit K::operator R(); since C++11
|
N/A | |
Note: for user-defined conversions, the return type implicitly and necessarily matches the operator name unless the type is inferred (e.g. operator auto(), operator decltype(auto)() etc.).
| ||||||
| dynamic cast conversion | dynamic_cast<R>(a) |
No | No | N/A | ||
| const cast conversion | const_cast<R>(a) |
No | No | N/A | ||
| reinterpret cast conversion | reinterpret_cast<R>(a) |
No | No | N/A | ||
| Allocate storage | new R[lower-alpha 14] |
Yes | No | void* K::operator new(size_t x);
|
void* operator new(size_t x);
| |
| Allocate storage (array) | new R[n]new R<:n:>new R??(n??)[lower-alpha 3][lower-alpha 4][lower-alpha 15] |
Yes | No | void* K::operator new[](size_t a);
|
void* operator new[](size_t a);
| |
| Deallocate storage | delete a |
Yes | No | void K::operator delete(void* a);
|
void operator delete(void* a);
| |
| Deallocate storage (array) | delete[] adelete<::> adelete??(??) a[lower-alpha 3][lower-alpha 4] |
Yes | No | void K::operator delete[](void* a);
|
void operator delete[](void* a);
| |
| Exception check since C++11 |
noexcept(a) |
No | No | N/A | ||
Notes:
- ↑ since C99
- ↑ Due to operator precedence ("." being higher than "*"),
*ptee.yis not correct; it is parsed as*(ptee.y)and thus the parentheses are necessary. - ↑ 3.0 3.1 3.2 Trigraphs were removed in C++17. They are still available in C as of C17 but will be removed in C23.
- ↑ 4.0 4.1 4.2 The brackets do not need to match as the trigraph bracket is substituted by the preprocessor and the digraph bracket is an alternative token that is equivalent. Only the cases where the brackets match are included since the other forms can be easily derived from the provided ones.
- ↑ Requires
iso646.hin C. See C++ operator synonyms - ↑ This alternative form is a side effect of the bitwise and alternative form for reasons explained in C++ operator synonyms
- ↑ The actual address of an object with an overloaded
operator &can be obtained withstd::addressof - ↑ The return type of
operator->()must be a type for which the->operation can be applied, such as a pointer type. Ifxis of typeCwhereCoverloadsoperator->(),x->ygets expanded tox.operator->()->y. - ↑ Meyers, Scott (Oct 1999), "Implementing operator->* for Smart Pointers", Dr. Dobb's Journal (Aristeia), http://aristeia.com/Papers/DDJ_Oct_1999.pdf.
- ↑ About C++11 User-defined literals
- ↑ The parentheses are not necessary when taking the size of a value, only when taking the size of a type. However, they are usually used regardless.
- ↑ C++ defines
alignofoperator, whereas C defines_Alignof(C23 defines both). Both operators have the same semantics. - ↑ As of C++23,
typeofandtypeof_unqualare not a part of the language. - ↑ The type name can also be inferred (e.g
new auto) if an initializer is provided. - ↑ The array size can also be inferred if an initializer is provided.
Cite error: <ref> tag with name "modulo" defined in <references> is not used in prior text.
Cite error: <ref> tag with name "bitshift" defined in <references> is not used in prior text.
Cite error: <ref> tag with name "trigraphs2" defined in <references> is not used in prior text.
<ref> tag with name "threewaycomparison" defined in <references> is not used in prior text.Operator precedence
The following is a table that lists the precedence and associativity of all the operators in the C and C++ languages. Operators are listed top to bottom, in descending precedence. Descending precedence refers to the priority of the grouping of operators and operands. Considering an expression, an operator which is listed on some row will be grouped prior to any operator that is listed on a row further below it. Operators that are in the same cell (there may be several rows of operators listed in a cell) are grouped with the same precedence, in the given direction. An operator's precedence is unaffected by overloading.
The syntax of expressions in C and C++ is specified by a phrase structure grammar.[14] The table given here has been inferred from the grammar.[citation needed] For the ISO C 1999 standard, section 6.5.6 note 71 states that the C grammar provided by the specification defines the precedence of the C operators, and also states that the operator precedence resulting from the grammar closely follows the specification's section ordering:
"The [C] syntax [i.e., grammar] specifies the precedence of operators in the evaluation of an expression, which is the same as the order of the major subclauses of this subclause, highest precedence first."[15]
A precedence table, while mostly adequate, cannot resolve a few details. In particular, note that the ternary operator allows any arbitrary expression as its middle operand, despite being listed as having higher precedence than the assignment and comma operators. Thus a ? b, c : d is interpreted as a ? (b, c) : d, and not as the meaningless (a ? b), (c : d). So, the expression in the middle of the conditional operator (between ? and :) is parsed as if parenthesized. Also, note that the immediate, unparenthesized result of a C cast expression cannot be the operand of sizeof. Therefore, sizeof (int) * x is interpreted as (sizeof(int)) * x and not sizeof ((int) * x).
Notes
The precedence table determines the order of binding in chained expressions, when it is not expressly specified by parentheses.
- For example,
++x*3is ambiguous without some precedence rule(s). The precedence table tells us that: x is 'bound' more tightly to ++ than to *, so that whatever ++ does (now or later—see below), it does it ONLY to x (and not tox*3); it is equivalent to (++x,x*3). - Similarly, with
3*x++, where though the post-fix ++ is designed to act AFTER the entire expression is evaluated, the precedence table makes it clear that ONLY x gets incremented (and NOT3*x). In fact, the expression (tmp=x++,3*tmp) is evaluated with tmp being a temporary value. It is functionally equivalent to something like (tmp=3*x,++x,tmp). - Abstracting the issue of precedence or binding, consider the diagram above for the expression 3+2*y[i]++. The compiler's job is to resolve the diagram into an expression, one in which several unary operators (call them 3+( . ), 2*( . ), ( . )++ and ( . )[ i ]) are competing to bind to y. The order of precedence table resolves the final sub-expression they each act upon: ( . )[ i ] acts only on y, ( . )++ acts only on y[i], 2*( . ) acts only on y[i]++ and 3+( . ) acts 'only' on 2*((y[i])++). It is important to note that WHAT sub-expression gets acted on by each operator is clear from the precedence table but WHEN each operator acts is not resolved by the precedence table; in this example, the ( . )++ operator acts only on y[i] by the precedence rules but binding levels alone do not indicate the timing of the postfix ++ (the ( . )++ operator acts only after y[i] is evaluated in the expression).
Many of the operators containing multi-character sequences are given "names" built from the operator name of each character. For example, += and -= are often called plus equal(s) and minus equal(s), instead of the more verbose "assignment by addition" and "assignment by subtraction".
The binding of operators in C and C++ is specified (in the corresponding Standards) by a factored language grammar, rather than a precedence table. This creates some subtle conflicts. For example, in C, the syntax for a conditional expression is:
logical-OR-expression ? expression : conditional-expressionwhile in C++ it is:
logical-OR-expression ? expression : assignment-expressionHence, the expression:
e = a < d ? a++ : a = dis parsed differently in the two languages. In C, this expression is a syntax error, because the syntax for an assignment expression in C is:
unary-expression '=' assignment-expressionIn C++, it is parsed as:
e = (a < d ? a++ : (a = d))which is a valid expression.[19][20]
If you want to use comma-as-operator within a single function argument, variable assignment, or other comma-separated list, you need to use parentheses,[21][22] e.g.:int a = 1, b = 2, weirdVariable = (++a, b), d = 4;Criticism of bitwise and equality operators precedence
The precedence of the bitwise logical operators has been criticized.[23] Conceptually, & and | are arithmetic operators like * and +.
The expression a & b == 7 is syntactically parsed as a & (b == 7) whereas the expression a + b == 7 is parsed as (a + b) == 7. This requires parentheses to be used more often than they otherwise would.
Historically, there was no syntactic distinction between the bitwise and logical operators. In BCPL, B and early C, the operators && || didn't exist. Instead & | had different meaning depending on whether they are used in a 'truth-value context' (i.e. when a Boolean value was expected, for example in if (a==b & c) {...} it behaved as a logical operator, but in c = a & b it behaved as a bitwise one). It was retained so as to keep backward compatibility with existing installations.[24]
Moreover, in C++ (and later versions of C) equality operations, with the exception of the three-way comparison operator, yield bool type values which are conceptually a single bit (1 or 0) and as such do not properly belong in "bitwise" operations.
C++ operator synonyms
C++ defines[25] certain keywords to act as aliases for a number of operators:
| Keyword | Operator |
|---|---|
and |
&&
|
and_eq |
&=
|
bitand |
&
|
bitor |
|
|
compl |
~
|
not |
!
|
not_eq |
!=
|
or |
||
|
or_eq |
|=
|
xor |
^
|
xor_eq |
^=
|
These can be used exactly the same way as the punctuation symbols they replace, as they are not the same operator under a different name, but rather simple token replacements for the name (character string) of the respective operator. This means that the expressions (a > 0 and not flag) and (a > 0 && !flag) have identical meanings. It also means that, for example, the bitand keyword may be used to replace not only the bitwise-and operator but also the address-of operator, and it can even be used to specify reference types (e.g., int bitand ref = n). The ISO C specification makes allowance for these keywords as preprocessor macros in the header file iso646.h. For compatibility with C, C++ provides the header ciso646, the inclusion of which has no effect.
See also
- Bitwise operations in C
- Bit manipulation
- Logical operator
- Boolean algebra (logic)
- Table of logic symbols
- Digraphs and trigraphs in C and in C++
References
- ↑ "C | Definition, History, Applications, & Facts | Britannica" (in en). Encyclopedia Britannica. https://www.britannica.com/technology/C-computer-programming-language.
- ↑ 2.0 2.1 2.2 Cite error: Invalid
<ref>tag; no text was provided for refs namedbk21st - ↑ "WG14-N2412: Two's complement sign representation". August 11, 2019. https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2412.pdf.
- ↑ "WG14-N2341: ISO/IEC TS 18661-2 - Floating-point extensions for C - Part 2: Decimal floating-point arithmetic". February 26, 2019. https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2341.pdf.
- ↑ "Unscrambling Declarations in C" (PDF). https://ptgmedia.pearsoncmg.com/images/0131774298/samplechapter/0131774298.pdf.
- ↑ Balagurusamy, E. Programming in ANSI C. Tata McGraw Hill. pp. 366.
- ↑ "The C Preprocessor: Implementation-defined behavior". https://gcc.gnu.org/onlinedocs/cpp/Implementation-defined-behavior.html.
- ↑ "String and character literals (C++)" (in en-us). https://docs.microsoft.com/en-us/cpp/cpp/string-and-character-literals-cpp?view=vs-2019#code-try-2.
- ↑ Kernighan & Richie
- ↑ "Operator overloading§Comparison operators". https://en.cppreference.com/w/cpp/language/operators#Comparison_operators.
- ↑ "Standard C++". https://isocpp.org/wiki/faq/operator-overloading.
- ↑ "user-defined conversion". https://en.cppreference.com/w/cpp/language/cast_operator.
- ↑ Explicit type conversion in C++
- ↑ ISO/IEC 9899:201x Programming Languages - C. open-std.org – The C Standards Committee. 19 December 2011. pp. 465.
- ↑ the ISO C 1999 standard, section 6.5.6 note 71 (Technical report). ISO. 1999.
- ↑ "C Operator Precedence - cppreference.com". https://en.cppreference.com/w/c/language/operator_precedence.
- ↑ "C++ Built-in Operators, Precedence and Associativity" (in en-us). https://docs.microsoft.com/en-US/cpp/cpp/cpp-built-in-operators-precedence-and-associativity.
- ↑ "C++ Operator Precedence - cppreference.com". https://en.cppreference.com/w/cpp/language/operator_precedence.
- ↑ "C Operator Precedence - cppreference.com". https://en.cppreference.com/w/c/language/operator_precedence.
- ↑ "Does the C/C++ ternary operator actually have the same precedence as assignment operators?". https://stackoverflow.com/questions/13515434/does-the-c-c-ternary-operator-actually-have-the-same-precedence-as-assignment/13515505.
- ↑ "Other operators - cppreference.com". https://en.cppreference.com/w/c/language/operator_other.
- ↑ "c++ - How does the Comma Operator work". https://stackoverflow.com/questions/54142/how-does-the-comma-operator-work/.
- ↑ C history § Neonatal C, Bell labs, https://www.bell-labs.com/usr/dmr/www/chist.html.
- ↑ "Re^10: next unless condition". https://www.perlmonks.org/?node_id=1159769.
- ↑ ISO/IEC 14882:1998(E) Programming Language C++. open-std.org – The C++ Standards Committee. 1 September 1998. pp. 40–41.
External links
- "Operators", C++ reference, https://en.cppreference.com/w/cpp/language/expressions#Operators.
- C Operator Precedence
- Postfix Increment and Decrement Operators: ++ and --, Microsoft, 17 August 2021, https://docs.microsoft.com/en-US/cpp/cpp/postfix-increment-and-decrement-operators-increment-and-decrement.
Control flow
Compound statement
A compound statement, a.k.a. statement block, is a matched pair of curly braces with any number of statements between, like:
{
{statement}
}
As a compound statement is a type of statement, when statement is used, it could be a single statement (without enclosing braces) or a compound statement (with enclosing braces). A compound statement is required for a function body and for a control structure branch that is not one statement.
A variable declared in a block can be referenced by code in that block (and inner blocks) below the declaration. Access to memory used for a block-declared variable after the block close (i.e. via a pointer) results in undefined behavior.
If statement
The if conditional statement is like:
if (expression) {
statement
} else {
statement
}
If the expression is not zero, control passes to the first statement. Otherwise, control passes to the second statement. If the else part is absent, then when the expression evaluates to zero, the first statement is simply skipped. An else always matches the nearest previous unmatched if. Braces may be used to override this when necessary, or for clarity.
Notably, the second statement can be another if statement. For example:
if (i == 1) {
printf("it's 1");
} else if (i == 2) {
printf("it's 2");
} else {
printf("it's something else");
}Switch statement
The switch transfers control to the case label that has a value matching the integral type expression or otherwise to the default label (if any). Control continues to the statements that follow the case label until a break statement or until the end of the switch statement. The syntax is like:
switch (expression) {
case label-name:
statement
case label-name:
statement
.
.
.
default:
statement
}
Each case value must be unique within the statement. There may be at most one default label.
Execution continues from one case label to the next if no break statement is encountered. This is known as falling through which is useful in some circumstances, but often is not not desired.
It is possible, although unusual, to locate case labels in the sub-blocks of inner control structures. Examples include Duff's device and Simon Tatham's implementation of coroutines in Putty.[1]
The following is an example of a switch over an int:
#include <stdio.h>
// ...
int num = 2;
switch (num) {
case 1:
printf("Number is 1\n");
break
case 2:
printf("Number is 2\n");
break;
case 3:
printf("Number is 3\n");
break;
default:
printf("Number is not 1, 2, or 3\n");
}As of C2Y, it is possible use a "case range" between two integer constants, using an ... (ellipsis).[2] There must be a space between the value and the ellipsis. This range is inclusive. For example:
#include <stdio.h>
// ...
int num = 2;
// new-style
switch (num) {
case 1 ... 3:
printf("Number is 1, 2, or 3\n");
break
default:
printf("Number is not 1, 2, or 3\n");
}
// old-style, using case fallthrough
switch (num) {
case 1:
case 2:
case 3:
printf("Number is 1, 2, or 3\n");
break
default:
printf("Number is not 1, 2, or 3\n");
}Iteration statement
There are three forms of iteration statement:
while (expression) {
statement
}
do {
statement
} while (expression)
for (init; test; next) {
statement
}
For the while and do statements, the sub-statement is executed repeatedly so long as the value of the expression is non-zero. For while, the test, including any side effects, occurs before each iteration. For do, the test occurs after each iteration. Thus, a do statement always executes its sub-statement at least once, whereas while might not execute the sub-statement at all.
The logic of for can be described in terms of while in that this:
for (e1; e2; e3) {
s;
}is equivalent to:
e1;
while (e2) {
s;
cont:
e3;
}except for the behavior of a continue; statement (which in the for loop jumps to e3 instead of e2). If e2 is blank, it would have to be replaced with a 1.
Any of the three expressions in the for loop may be omitted. A missing second expression makes the while test always non-zero; describing an infinite loop.
Since C99, the first expression may take the form of a declaration with scope limited to the sub-statement For example:
for (int i = 0; i < limit; ++i) {
// ...
}The foreach loop does not exist in C, like it does in Java and C++. However, it can be emulated using macros.
Jump statement
There are four jump statements (transfer control unconditionally): goto, continue, break, and return.
The goto statement passes program control to a labeled statement. It has the following syntax:
goto label-name
A continue statement which is simply the word continue, transfers control to the loop-continuation point of the innermost, enclosing iteration statement. It must be enclosed within an iteration statement. For example:
while (true) {
// ...
continue;
}
do {
// ...
continue;
} while (true);
for (;;) {
// ...
continue;
}The break statement which is simply the word break ends a for, while, do, or switch statement. Control passes to the statement following the enclosing control statement.
The return statement transfers control the function caller. When return is followed by an expression, the value is returned to the caller. Encountering the end of the function is equivalent to a return with no expression. In that case, if the function is declared as returning a value and the caller tries to use the returned value, the behavior is undefined.
Labels
A label marks a point in the code to which control can be transferred. A label is an identifier followed by a colon. For example:
if (i == 1) {
goto END;
}
// other code
END:The standard does not define a method to retrieve the address of a label, but
GCC extends the language with a unary && operator that returns the address of a label. The address can be stored in a void* variable and may be used later with a goto. This feature can be used to implement a jump table.
For example, the following prints hi repeatedly:
void* ptr = &&J1;
J1: printf("hi");
goto *ptr;Since C2Y, C has labelled loops, similar to Java. This allows for attaching labels to for, and transferring control through break and continue with labels (multi-level breaks).[3]
// using break:
outer:
for (int i = 0; i < n; ++i) {
switch (i) {
case 1:
break; // jumps to 1
case 2:
break outer; // jumps to 2
default:
continue;
}
// 1
}
// 2
// using continue:
outer:
for (int i = 0; i < m; ++i) {
for (int j = 0; j < n; ++j) {
continue; // jumps to 1
continue outer; // jumps to 2
// 1
}
// 2
}Functions
Definition
For a function that returns a value, a definition consists of a return type name, a function name that is unique in the codebase, a list of parameters in parentheses, and a statement block that ends with a return statement. The block can contain a return statement to exit the function before the end of the block. The syntax is like:
type-name function-name(parameter-list)
{
statement-list
return value;
}
A function that returns no value is declared with void instead of a type name, like:
void function-name(parameter-list)
{
statement-list
}
The standard does not include lambda functions, but some translators do.
Parameters
A parameter-list is a comma-separated list of formal parameter declarations; each item a type name followed by a variable name:
type-name variable-name{, type-name variable-name}
The return type cannot be an array or a function. For example:
int f()[3]; // Error: function returning an array
int (*g())[3]; // OK: function returning a pointer to an array
void h()(); // Error: function returning a function
void (*k())(); // OK: function returning a function pointerIf the function accepts no parameters, the parameter-list may be the keyword void or blank, but these have different implications. Calling a function with arguments when it is declared with void for the parameter-list is invalid syntax. Calling a function with arguments when it is declared with a blank parameter-list is not invalid syntax, but may result in undefined behavior. Using void, is therefore, best practice.
A function can accept a variable number of arguments by including ... at the end of the argument list. A commonly used function with this declaration is the standard library function printf which has prototype:
int printf(const char*, ...);Consuming variable length arguments can be accomplished via standard library functions declared in <stdarg.h>.
Calling
Code can access a function of a library if it is both declared and defined. Often a declaration is provided for a library function via a header file that the consuming code uses via the #include directive. Alternatively, the consuming code can declare the function in its own file. The function definition is associated with the consuming code at link-time. The standard library is generally linked by default whereas other libraries require link-time configuration.
Accessing a user-defined function that is defined in a different file is similar to using a library function. The consuming code declares the function either by including a header file or directly in its file. Linking to the definition in the other file is handled when the object files are linked.
Calling a function that is defined in the same file is relatively simple. The definition or a declaration of it must be above the call.
Argument passing
An argument is passed to a function by value which means that a called function receives a copy of the argument and cannot alter the argument variable. For a function to alter the value of a variable, the caller passes the variable's address (a pointer) which simulates what other languages provide as by reference. The called function can modify the variable by dereferencing the passed address.
In the following code, the address of x is passed by specifying &x in the call. The called function receives the address as y and accesses x as *y.
void incInt(int* y) {
(*y)++;
}
int main(void) {
int x = 7;
incInt(&x);
return 0;
}The following code demonstrates a more advanced use of pointers – passing a pointer to a pointer. An int pointer named a is defined on line 9 and its address is passed to the function on line 10. The function receives a pointer to pointer to int named a_p. It assigns a (as *a_p). After the call, on line 11, the memory allocated and assigned to address a is freed.
#include <stdio.h>
#include <stdlib.h>
void allocate_array(int** const p, const int count) {
*p = malloc(sizeof(int) * count);
}
int main(void) {
int* a;
allocate_array(&a, 42);
free(a);
return 0;
}Array passing
Function parameters of array type may at first glance appear to be an exception to the pass-by-value rule as demonstrated by the following program that prints 123, not 1:
#include <stdio.h>
void setArray(int a[], int index) {
array[index] = 123;
}
int main(void) {
int a[1] = {1};
setArray(a, 0);
printf("a[0]=%d\n", a[0]);
return 0;
}However, there is a different reason for this behavior. An array parameter is treated as a pointer. The following prototype is equivalent to the function prototype above:
void setArray(int* a, int index);At the same time, rules for the use of arrays in expressions cause the value of a to be treated as a pointer to the first element. Thus, this is still pass-by-value, with the caveat that it is the address of the first element of the array being passed by value, not the contents of the array.
Since C99, the programmer can specify that a function takes an array of a certain size by using the keyword static. In void setArray(int array[static 4], int index) the first parameter must be a pointer to the first element of an array of length at least 4. It is also possible to use qualifiers (const, volatile and restrict) to the pointer type that the array is converted to.
Attributes
Added in C23 and originating from C++11, C supports attribute specifier sequences.[4] Attributes can be applied to any symbol that supports them, including functions and variables, and any symbol marked with an attribute will be specifically treated by the compiler as necessary. These can be thought of as similar to Java annotations for providing additional information to the compiler, however they differ in that attributes in C are not metadata that is meant to be accessed using reflection. Furthermore, one cannot create custom attributes in C, unlike in Java where one may define custom annotations in addition to the standard ones. However, C does have implementation/vendor-specific attributes which are non-standard. These typically have a namespace associated with them. For instance, GCC and Clang have attributes under the gnu:: namespace, and all such attributes are of the form gnu::*, though C does not have support for namespacing in the language.
The syntax of using an attribute on a function is like so:
nodiscard
bool satisfiesProperty(const struct MyStruct* s);The standard defines the following attributes:
| Name | Description |
|---|---|
noreturn
|
Indicates that the specified function will not return to its caller. |
deprecateddeprecated("reason")
|
Indicates that the use of the marked symbol is allowed but discouraged/deprecated for the reason specified (if given). |
fallthrough
|
Indicates that the fall through from the previous case label is intentional. |
maybe unused
|
Suppresses compiler warnings on an unused entity. |
nodiscardnodiscard("reason")
|
Issues a compiler warning if the return value of the marked symbol is discarded or ignored for the reason specified (if given). |
unsequenced
|
Indicates that a function is stateless, effectless, idempotent and independent. |
reproducible
|
Indicates that a function is effectless and idempotent. |
Dynamic memory
C dynamic memory allocation refers to performing manual memory management for dynamic memory allocation in the C programming language via a group of functions in the C standard library, mainly malloc, realloc, calloc, aligned_alloc and free.[5][6][7]
The C++ programming language includes these functions; however, the operators new and delete provide similar functionality and are recommended by that language's authors.[8] Still, there are several situations in which using new/delete is not applicable, such as garbage collection code or performance-sensitive code, and a combination of malloc and new may be required instead of the higher-level new operator.
Many different implementations of the actual memory allocation mechanism, used by malloc, are available. Their performance varies in both execution time and required memory.
Rationale
The C programming language manages memory statically, automatically, or dynamically. Static-duration variables are allocated in main memory, usually along with the executable code of the program, and persist for the lifetime of the program; automatic-duration variables are allocated on the stack and come and go as functions are called and return. For static-duration and automatic-duration variables, the size of the allocation must be compile-time constant (except for the case of variable-length automatic arrays[9]). If the required size is not known until run-time (for example, if data of arbitrary size is being read from the user or from a disk file), then using fixed-size data objects is inadequate.
The lifetime of allocated memory can also cause concern. Neither static- nor automatic-duration memory is adequate for all situations. Automatic-allocated data cannot persist across multiple function calls, while static data persists for the life of the program whether it is needed or not. In many situations the programmer requires greater flexibility in managing the lifetime of allocated memory.
These limitations are avoided by using dynamic memory allocation, in which memory is more explicitly (but more flexibly) managed, typically by allocating it from the heap (free storage), an area of memory structured for this purpose. In C, the library function malloc is used to allocate a block of memory on the heap. The program accesses this block of memory via a pointer that malloc returns. When the memory is no longer needed, the pointer is passed to free which deallocates the memory so that it can be used for other purposes.
The original description of C indicated that calloc and cfree were in the standard library, but not malloc. Code for a simple model implementation of a storage manager for Unix was given with alloc and free as the user interface functions, and using the sbrk system call to request memory from the operating system.[10] The 6th Edition Unix documentation gives alloc and free as the low-level memory allocation functions.[11] The malloc and free routines in their modern form are completely described in the 7th Edition Unix manual.[12][13]
Some platforms provide library or intrinsic function calls which allow run-time dynamic allocation from the C stack rather than the heap (e.g. alloca()[14]). This memory is automatically freed when the calling function ends.
Overview of functions
The C dynamic memory allocation functions are defined in <stdlib.h> header (<cstdlib> header in C++).[5]
Differences between malloc() and calloc()
malloc()takes a single argument (the amount of memory to allocate in bytes), whilecalloc()takes two arguments — the number of elements and the size of each element.malloc()only allocates memory, whilecalloc()allocates and sets the bytes in the allocated region to zero.[15]
Usage example
Creating an array of ten integers with automatic scope is straightforward in C:
int a[10];However, the size of the array is fixed at compile time. If one wishes to allocate a similar array dynamically without using a variable-length array, which is not guaranteed to be supported in all C11 implementations, an array can be allocated using malloc, which returns a void pointer (indicating that it is a pointer to a region of unknown data type), which can be cast for safety.
int* a = (int*)malloc(10 * sizeof(int));This computes the number of bytes that ten integers occupy in memory, then requests that many bytes from malloc and assigns the result to a pointer named a (due to C syntax, pointers and arrays can be used interchangeably in some situations).
Because malloc might not be able to service the request, it might return a null pointer and it is good programming practice to check for this:
int* a = (int*)malloc(10 * sizeof(int));
if (!a) {
fprintf(stderr, "malloc failed\n");
return -1;
}When the program no longer needs the dynamic array, it must eventually call free to return the memory it occupies to the free store:
free(a);The memory set aside by malloc is not initialized and may contain cruft: the remnants of previously used and discarded data. After allocation with malloc, elements of the array are uninitialized variables. The command calloc will return an allocation that has already been cleared:
int* a = (int*)calloc(10, sizeof(int));With realloc we can resize the amount of memory a pointer points to. For example, if we have a pointer acting as an array of size and we want to change it to an array of size , we can use realloc.
int* a = (int*)malloc(2 * sizeof(int));
a[0] = 1;
a[1] = 2;
a = (int*)realloc(a, 3 * sizeof(int));
a[2] = 3;Note that realloc must be assumed to have changed the base address of the block (i.e. if it has failed to extend the size of the original block, and has therefore allocated a new larger block elsewhere and copied the old contents into it). Therefore, any pointers to addresses within the original block are also no longer valid.
Type safety
As stated earlier, malloc returns a void pointer (void*), which indicates that it is a pointer to a region of unknown data type. The use of casting is required in C++ due to the strong type system, whereas this is not the case in C. One may "cast" (see type conversion) this pointer to a specific type:
// without a cast
int* ptr1 = malloc(10 * sizeof(*ptr));
// with a cast
int* ptr2 = (int*)malloc(10 * sizeof(*ptr));There are advantages and disadvantages to performing such a cast. Including the cast improves interoperability between C and C++, allowing C code to be compiled or used as C++. Furthermore, the cast allows for pre-1989 versions of malloc that originally returned a char*.[16] Additionally, casting can help the developer identify inconsistencies in type sizing should the destination pointer type change, particularly if the pointer is declared far from the malloc() call (although modern compilers and static analysers can warn on such behaviour without requiring the cast[17]).
However, under the C standard, the cast is not necessary, and adding the cast may mask failure to include the header <stdlib.h>, in which the function prototype for malloc is found.[16][18] In the absence of a prototype for malloc, the C90 standard requires that the C compiler assume malloc returns an int. If there is no cast, C90 requires a diagnostic when this integer is assigned to the pointer; however, with the cast, this diagnostic would not be produced, hiding a bug. On certain architectures and data models (such as LP64 on 64-bit systems, where long and pointers are 64-bit and int is 32-bit), this error can actually result in undefined behaviour, as the implicitly declared malloc returns a 32-bit value whereas the actually defined function returns a 64-bit value. Depending on calling conventions and memory layout, this may result in stack smashing. This issue is less likely to go unnoticed in modern compilers, as C99 does not permit implicit declarations, so the compiler must produce a diagnostic even if it does assume int return. Additionally, if the type of the pointer is changed at its declaration, the lines where malloc is called and the cast will have to be updated as well.
In C++, if std::malloc must be used, it is preferable to static_cast it, rather than use a raw cast:
int* p = static_cast<int*>(malloc(10 * sizeof(*ptr)));Common errors
The improper use of dynamic memory allocation can frequently be a source of bugs. These can include security bugs or program crashes, most often due to segmentation faults.
Most common errors are as follows:[19]
- Not checking for allocation failures
- Memory allocation is not guaranteed to succeed, and may instead return a null pointer. Using the returned value, without checking if the allocation is successful, invokes undefined behavior. This usually leads to crash (due to the resulting segmentation fault on the null pointer dereference), but there is no guarantee that a crash will happen so relying on that can also lead to problems.
- Memory leaks
- Failure to deallocate memory using
freeleads to the buildup of non-reusable memory, which is no longer used by the program. This wastes memory resources and can lead to allocation failures when these resources are exhausted. - Logical errors
- All allocations must follow the same pattern: allocation using
malloc, usage to store data, deallocation usingfree. Failures to adhere to this pattern, such as memory usage after a call tofree(dangling pointer) or before a call tomalloc(wild pointer), callingfreetwice ("double free"), etc., usually causes a segmentation fault and results in a crash of the program. These errors can be transient and hard to debug – for example, freed memory is usually not immediately reclaimed by the OS, and thus dangling pointers may persist for a while and appear to work.
In addition, as an interface that precedes ANSI C standardization, malloc and its associated functions have behaviors that were intentionally left to the implementation to define for themselves. One of them is the zero-length allocation, which is more of a problem with realloc since it is more common to resize to zero.[20] Although both POSIX and the Single Unix Specification require proper handling of 0-size allocations by either returning NULL or something else that can be safely freed,[21] not all platforms are required to abide by these rules. Among the many double-free errors that it has led to, the 2019 WhatsApp RCE was especially prominent.[22] A way to wrap these functions to make them safer is by simply checking for 0-size allocations and turning them into those of size 1. (Returning NULL has its own problems: it otherwise indicates an out-of-memory failure. In the case of realloc it would have signaled that the original memory was not moved and freed, which again is not the case for size 0, leading to the double-free.)[23]
Implementations
The implementation of memory management depends greatly upon operating system and architecture. Some operating systems supply an allocator for malloc, while others supply functions to control certain regions of data. The same dynamic memory allocator is often used to implement both malloc and the operator new in C++.[24]
Heap-based
Implementation of legacy allocators was commonly done using the heap segment. The allocator would usually expand and contract the heap to fulfill allocation requests.
The heap method suffers from a few inherent flaws:
- A linear allocator can only shrink if the last allocation is released. Even if largely unused, the heap can get "stuck" at a very large size because of a small but long-lived allocation at its tip which could waste any amount of address space, although some allocators on some systems may be able to release entirely empty intermediate pages to the OS.
- A linear allocator is sensitive to fragmentation. A good allocator will attempt to track and reuse free slots through the entire heap, but as allocation sizes and lifetimes get mixed it can be difficult and expensive to find or coalesce free segments large enough to hold new allocation requests.
- A linear allocator has extremely poor concurrency characteristics, as the heap segment is per-process every thread has to synchronise on allocation, and concurrent allocations from threads which may have very different work loads amplifies the previous two issues.
dlmalloc and ptmalloc
Doug Lea has developed the public domain dlmalloc ("Doug Lea's Malloc") as a general-purpose allocator, starting in 1987. The GNU C library (glibc) is derived from Wolfram Gloger's ptmalloc ("pthreads malloc"), a fork of dlmalloc with threading-related improvements.[25][26][27] As of November 2023, the latest version of dlmalloc is version 2.8.6 from August 2012. In 2023 it was relicensed from CC0 to MIT-0.[28]
dlmalloc is a boundary tag allocator. Memory on the heap is allocated as "chunks", an 8-byte aligned data structure which contains a header, and usable memory. Allocated memory contains an 8- or 16-byte overhead for the size of the chunk and usage flags (similar to a dope vector). Unallocated chunks also store pointers to other free chunks in the usable space area, making the minimum chunk size 16 bytes on 32-bit systems and 24/32 (depends on alignment) bytes on 64-bit systems.[26][28]: 2.8.6, Minimum allocated size
Unallocated memory is grouped into "bins" of similar sizes, implemented by using a double-linked list of chunks (with pointers stored in the unallocated space inside the chunk). Bins are sorted by size into three classes:[26][28]: Overlaid data structures
- For requests below 256 bytes (a "smallbin" request), a simple two power best fit allocator is used. If there are no free blocks in that bin, a block from the next highest bin is split in two.
- For requests of 256 bytes or above but below the mmap threshold, dlmalloc since v2.8.0 use an in-place bitwise trie algorithm ("treebin"). If there is no free space left to satisfy the request, dlmalloc tries to increase the size of the heap, usually via the brk system call. This feature was introduced way after ptmalloc was created (from v2.7.x), and as a result is not a part of glibc, which inherits the old best-fit allocator.
- For requests above the mmap threshold (a "largebin" request), the memory is always allocated using the mmap system call. The threshold is usually 128 KB.[29] The mmap method averts problems with huge buffers trapping a small allocation at the end after their expiration, but always allocates an entire page of memory, which on many architectures is 4096 bytes in size.[30]
Game developer Adrian Stone argues that dlmalloc, as a boundary-tag allocator, is unfriendly for console systems that have virtual memory but do not have demand paging. This is because its pool-shrinking and growing callbacks (sysmalloc/systrim) cannot be used to allocate and commit individual pages of virtual memory. In the absence of demand paging, fragmentation becomes a greater concern.[31]
FreeBSD's and NetBSD's jemalloc
Since FreeBSD 7.0 and NetBSD 5.0, the old malloc implementation (phkmalloc by Poul-Henning Kamp) was replaced by jemalloc, written by Jason Evans. The main reason for this was a lack of scalability of phkmalloc in terms of multithreading. In order to avoid lock contention, jemalloc uses separate "arenas" for each CPU. Experiments measuring number of allocations per second in multithreading application have shown that this makes it scale linearly with the number of threads, while for both phkmalloc and dlmalloc performance was inversely proportional to the number of threads.[32]
OpenBSD's malloc
OpenBSD's implementation of the malloc function makes use of mmap. For requests greater in size than one page, the entire allocation is retrieved using mmap; smaller sizes are assigned from memory pools maintained by malloc within a number of "bucket pages", also allocated with mmap.[33] On a call to free, memory is released and unmapped from the process address space using munmap. This system is designed to improve security by taking advantage of the address space layout randomization and gap page features implemented as part of OpenBSD's mmap system call, and to detect use-after-free bugs—as a large memory allocation is completely unmapped after it is freed, further use causes a segmentation fault and termination of the program.
The GrapheneOS project initially started out by porting OpenBSD's memory allocator to Android's Bionic C Library.[34]
Hoard malloc
Hoard is an allocator whose goal is scalable memory allocation performance. Like OpenBSD's allocator, Hoard uses mmap exclusively, but manages memory in chunks of 64 kilobytes called superblocks. Hoard's heap is logically divided into a single global heap and a number of per-processor heaps. In addition, there is a thread-local cache that can hold a limited number of superblocks. By allocating only from superblocks on the local per-thread or per-processor heap, and moving mostly-empty superblocks to the global heap so they can be reused by other processors, Hoard keeps fragmentation low while achieving near linear scalability with the number of threads.[35]
mimalloc
An open-source compact general-purpose memory allocator from Microsoft Research with focus on performance.[36] The library is about 11,000 lines of code.
Thread-caching malloc (tcmalloc)
Every thread has a thread-local storage for small allocations. For large allocations mmap or sbrk can be used. TCMalloc, a malloc developed by Google,[37] has garbage-collection for local storage of dead threads. The TCMalloc is considered to be more than twice as fast as glibc's ptmalloc for multithreaded programs.[38][39]
In-kernel
Operating system kernels need to allocate memory just as application programs do. The implementation of malloc within a kernel often differs significantly from the implementations used by C libraries, however. For example, memory buffers might need to conform to special restrictions imposed by DMA, or the memory allocation function might be called from interrupt context.[40] This necessitates a malloc implementation tightly integrated with the virtual memory subsystem of the operating system kernel.
Overriding malloc
Because malloc and its relatives can have a strong impact on the performance of a program, it is not uncommon to override the functions for a specific application by custom implementations that are optimized for application's allocation patterns. The C standard does not specify a way of doing this, but operating systems have found various ways to do this by exploiting dynamic linking. One way is to simply link in a different library to override the symbols. Another, employed by Unix System V.3, is to make malloc and free function pointers that an application can reset to custom functions.[41]
In the source code the function can easily be replaced by using preprocessor macros, such as
#define malloc custom_malloc #define realloc custom_realloc #define free custom_free
which get defined in a header that all files that use memory allocation functions must include.
Theese macros should also replace the function strdup, since its result is passed to free too. Data allocated inside of library functions (like FILE structures in the stdio.h) still get directly handled without using the wrappers.
The most common form on POSIX-like systems is to set the environment variable LD_PRELOAD with the path of the allocator, so that the dynamic linker uses that version of malloc/calloc/free instead of the libc implementation.
Allocation size limits
The largest possible memory block malloc can allocate depends on the host system, particularly the size of physical memory and the operating system implementation.
Theoretically, the largest number should be the maximum value that can be held in a size t type, which is an implementation-dependent unsigned integer representing the size of an area of memory. In the C99 standard and later, it is available as the SIZE_MAX constant from <stdint.h>. Although not guaranteed by ISO C, it is usually 2^(CHAR_BIT * sizeof(size_t)) - 1.
On glibc systems, the largest possible memory block malloc can allocate is only half this size, namely 2^(CHAR_BIT * sizeof(ptrdiff_t) - 1) - 1.[42]
On systems that use segmentation, like some memory models under MS-DOS, the largest possible memory block is typically one segment or less and thus significant smaller than the total adress space, and size_t might be a smaller type than ptrdiff_t.
Extensions and alternatives
The C library implementations shipping with various operating systems and compilers may come with alternatives and extensions to the standard malloc interface. Notable among these is:
alloca, which allocates a requested number of bytes on the call stack. No corresponding deallocation function exists, as typically the memory is deallocated as soon as the calling function returns.allocawas present on Unix systems as early as 32/V (1978), but its use can be problematic in some (e.g., embedded) contexts.[43] While supported by many compilers, it is not part of the ANSI-C standard and therefore may not always be portable. It may also cause minor performance problems: it leads to variable-size stack frames, so that both stack and frame pointers need to be managed (with fixed-size stack frames, one of these is redundant).[44] Larger allocations may also increase the risk of undefined behavior due to a stack overflow.[45] C99 offered variable-length arrays as an alternative stack allocation mechanism – however, this feature was relegated to optional in the later C11 standard.- POSIX defines a function
posix_memalignthat allocates memory with caller-specified alignment. Its allocations are deallocated withfree,[46] so the implementation usually needs to be a part of the malloc library.
See also
References
- ↑ Tatham, Simon (2000). "Coroutines in C". https://www.chiark.greenend.org.uk/~sgtatham/coroutines.html.
- ↑ "WG14-N3370: Case range expressions, v3.1". October 1, 2024. https://www.open-std.org/JTC1/SC22/WG14/www/docs/n3370.htm.
- ↑ Alex Celeste (18 September 2024). "Named loops, v3". WG 14. https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3355.htm.
- ↑ "Attribute specifier sequence (since C23)". https://en.cppreference.com/w/c/language/attributes.html.
- ↑ 5.0 5.1 7.20.3 Memory management functions (PDF). ISO/IEC 9899:1999 specification (Technical report). p. 313.
- ↑ Summit, Steve. "Chapter 11: Memory Allocation". C Programming Notes. http://www.eskimo.com/~scs/cclass/notes/sx11.html.
- ↑ "aligned_alloc(3) - Linux man page". https://linux.die.net/man/3/aligned_alloc.
- ↑ Stroustrup, Bjarne (2008). Programming: Principles and Practice Using C++. Addison Wesley. p. 1009. ISBN 978-0-321-54372-1.
- ↑ "gcc manual". gnu.org. https://gcc.gnu.org/onlinedocs/gcc/Variable-Length.html.
- ↑ Brian W. Kernighan, Dennis M. Ritchie, The C Programming Language, Prentice-Hall, 1978; Section 7.9 (page 156) describes
callocandcfree, and Section 8.7 (page 173) describes an implementation forallocandfree. - ↑ – Template:Man/v6
- ↑ – Version 7 Unix Programmer's Manual
- ↑ Anonymous, Unix Programmer's Manual, Vol. 1, Holt Rinehart and Winston, 1983 (copyright held by Bell Telephone Laboratories, 1983, 1979); The
manpage formallocetc. is given on page 275. - ↑ – FreeBSD Library Functions Manual
- ↑ – Linux Programmer's Manual – Library Functions
- ↑ 16.0 16.1 "Casting malloc". Cprogramming.com. http://faq.cprogramming.com/cgi-bin/smartfaq.cgi?id=1043284351&answer=1047673478.
- ↑ "clang: lib/StaticAnalyzer/Checkers/MallocSizeofChecker.cpp Source File". http://clang.llvm.org/doxygen/MallocSizeofChecker_8cpp_source.html.
- ↑ "comp.lang.c FAQ list · Question 7.7b". C-FAQ. http://www.c-faq.com/malloc/mallocnocast.html.
- ↑ Reek, Kenneth (1997-08-04) (in en). Pointers on C (1 ed.). Pearson. ISBN 9780673999863.
- ↑ "MEM04-C. Beware of zero-length allocations - SEI CERT C Coding Standard - Confluence". https://wiki.sei.cmu.edu/confluence/display/c/MEM04-C.+Beware+of+zero-length+allocations.
- ↑ "POSIX.1-2017: malloc". https://pubs.opengroup.org/onlinepubs/9699919799/.
- ↑ Awakened (2 October 2019). "How a double-free bug in WhatsApp turns to RCE". https://awakened1712.github.io/hacking/hacking-whatsapp-gif-rce/.
- ↑ @RichFelker (3 October 2019). "Wow. The WhatsApp RCE was the wrong behavior for realloc(p,0) so many implementations insist on.". https://twitter.com/RichFelker/status/1179701167569416192.
- ↑ Alexandrescu, Andrei (2001). Modern C++ Design: Generic Programming and Design Patterns Applied. Addison-Wesley. p. 78.
- ↑ "Wolfram Gloger's malloc homepage". http://malloc.de/en/.
- ↑ 26.0 26.1 26.2 Kaempf, Michel (2001). "Vudo malloc tricks". Phrack (57): 8. http://phrack.org/issues/57/8.html. Retrieved 29 April 2009.
- ↑ "Glibc: Malloc Internals". https://sourceware.org/glibc/wiki/MallocInternals.
- ↑ 28.0 28.1 28.2 Lee, Doug. "A Memory Allocator". http://gee.cs.oswego.edu/dl/html/malloc.html. HTTP for Source Code
- ↑ "Malloc Tunable Parameters". GNU. https://www.gnu.org/software/libc/manual/html_node/Malloc-Tunable-Parameters.html.
- ↑ Sanderson, Bruce (12 December 2004). "RAM, Virtual Memory, Pagefile and all that stuff". Microsoft Help and Support. http://support.microsoft.com/kb/555223.
- ↑ Stone, Adrian. "The Hole That dlmalloc Can't Fill". http://gameangst.com/?p=496.
- ↑ Evans, Jason (16 April 2006). "A Scalable Concurrent malloc(3) Implementation for FreeBSD". http://people.freebsd.org/~jasone/jemalloc/bsdcan2006/jemalloc.pdf.
- ↑ "libc/stdlib/malloc.c". http://bxr.su/OpenBSD/lib/libc/stdlib/malloc.c.
- ↑ "History | GrapheneOS" (in en). https://grapheneos.org/history/.
- ↑ Berger, E. D.; McKinley, K. S.; Blumofe, R. D.; Wilson, P. R. (November 2000). "Hoard: A Scalable Memory Allocator for Multithreaded Applications". ASPLOS-IX. pp. 117–128. doi:10.1145/378993.379232. ISBN 1-58113-317-0. http://www.cs.umass.edu/~emery/pubs/berger-asplos2000.pdf.
- ↑ "Microsoft releases optimized malloc() as open source - Slashdot". https://slashdot.org/firehose.pl?op=view&id=110928452.
- ↑ TCMalloc homepage
- ↑ Ghemawat, Sanjay; Menage, Paul; TCMalloc : Thread-Caching Malloc
- ↑ Callaghan, Mark (18 January 2009). "High Availability MySQL: Double sysbench throughput with TCMalloc". Mysqlha.blogspot.com. https://mysqlha.blogspot.com/2009/01/double-sysbench-throughput-with_18.html.
- ↑ "kmalloc()/kfree() include/linux/slab.h". People.netfilter.org. http://people.netfilter.org/rusty/unreliable-guides/kernel-hacking/routines-kmalloc.html.
- ↑ "Chapter 9: Shared libraries". Linkers and Loaders. The Morgan Kaufmann Series in Software Engineering and Programming (1 ed.). San Francisco, USA: Morgan Kaufmann. 2000. ISBN 1-55860-496-0. OCLC 42413382. https://archive.today/20130126001629/http://www.iecc.com/linker/linker09.html. Retrieved 2020-01-12. Code: [1][2] Errata: [3]
- ↑ "malloc: make malloc fail with requests larger than PTRDIFF_MAX". 18 April 2019. https://sourceware.org/bugzilla/show_bug.cgi?id=23741#c2.
- ↑ "Why is the use of alloca() not considered good practice?". https://stackoverflow.com/questions/1018853/why-is-the-use-of-alloca-not-considered-good-practice.
- ↑ Amarasinghe, Saman; Leiserson, Charles (2010). "6.172 Performance Engineering of Software Systems, Lecture 10". Massachusetts Institute of Technology. http://ocw.mit.edu/courses/electrical-engineering-and-computer-science/6-172-performance-engineering-of-software-systems-fall-2010/video-lectures/lecture-8-cache-efficient-algorithms/.
- ↑ "alloca(3) - Linux manual page". http://man7.org/linux/man-pages/man3/alloca.3.html.
- ↑ – System Interfaces Reference, The Single UNIX Specification, Issue 7 from The Open Group
External links
| The Wikibook C Programming has a page on the topic of: C Programming/C Reference |
- Definition of malloc in IEEE Std 1003.1 standard
- Lea, Doug; The design of the basis of the glibc allocator
- Gloger, Wolfram; The ptmalloc homepage
- Berger, Emery; The Hoard homepage
- Douglas, Niall; The nedmalloc homepage
- Evans, Jason; The jemalloc homepage
- Google; The tcmalloc homepage
- Simple Memory Allocation Algorithms on OSDEV Community
- Michael, Maged M.; Scalable Lock-Free Dynamic Memory Allocation
- Bartlett, Jonathan; Inside memory management – The choices, tradeoffs, and implementations of dynamic allocation
- Memory Reduction (GNOME) wiki page with much information about fixing malloc
- C99 standard draft, including TC1/TC2/TC3
- Some useful references about C
- ISO/IEC 9899 – Programming languages – C
- Understanding glibc malloc
See also
- Blocks (C language extension)
- C data types
- C standard library
- C++ syntax
- C# syntax
- Java syntax
- Rust syntax
- List of C-family programming languages
- Operators in C and C++
Notes
References
<ref> tag with name "bk21st" defined in <references> is not used in prior text.- General
- Kernighan, Brian W.; Ritchie, Dennis M. (1988). The C Programming Language (2nd ed.). Upper Saddle River, New Jersey: Prentice Hall PTR. ISBN 0-13-110370-9. https://archive.org/details/cprogramminglang00bria.
- American National Standard for Information Systems - Programming Language - C - ANSI X3.159-1989
- "ISO/IEC 9899:2018 - Information technology -- Programming languages -- C". https://www.iso.org/standard/74528.html.
- "ISO/IEC 9899:1999 - Programming languages - C". Iso.org. 2011-12-08. https://www.iso.org/iso/iso_catalogue/catalogue_ics/catalogue_detail_ics.htm?csnumber=29237.
External links
- The syntax of C in Backus–Naur form
- Programming in C
- The comp.lang.c Frequently Asked Questions Page
- C reference
