Extracting urls from @font-face by searching within @font-face for replacement
css, php, regex
Solution
You can use this:
$pattern = <<<'LOD'
~
(?(DEFINE)
(?<quoted_content>
(["']) (?>[^"'\\]++ | \\{2} | \\. | (?!\g{-1})["'] )*+ \g{-1}
)
(?<comment> /\* .*? \*/ )
(?<url_skip> (?: https?: | data: ) [^"'\s)}]*+ )
(?<other_content>
(?> [^u}/"']++ | \g<quoted_content> | \g<comment>
| \Bu | u(?!rl\s*+\() | /(?!\*)
| \g<url_start> \g<url_skip> ["']?+
)++
)
(?<anchor> \G(?<!^) ["']?+ | @font-face \s*+ { )
(?<url_start> url\( \s*+ ["']?+ )
)
\g<comment> (*SKIP)(*FAIL) |
\g<anchor> \g<other_content>?+ \g<url_start> \K [./]*+
( [^"'\s)}]*+ ) # url
~xs
LOD;
$result = preg_replace($pattern, 'http://cdn.test.com/fonts/$8', $data);
print_r($result);
test string
$data = <<<'LOD'
@font-face {
font-family: 'FontAwesome';
src: url("fonts/fontawesome-webfont.eot?v=4.0.3");
src: url(fonts/fontawesome-webfont.eot?#iefix&v=4.0.3) format("embedded-opentype"),
/*url("fonts/fontawesome-webfont.woff?v=4.0.3") format("woff"),*/
url("http://domain.com/fonts/fontawesome-webfont.ttf?v=4.0.3") format("truetype"),
url('fonts/fontawesome-webfont.svg?v=4.0.3#fontawesomeregular') format("svg");
font-weight: normal;
font-style: normal;
}
/*
@font-face {
font-family: 'Font1';
src: url("fonts/font1.eot");
} */
@font-face {
font-family: 'Fon\'t2';
src: url("fonts/font2.eot");
}
@font-face {
font-family: 'Font3';
src: url("../fonts/font3.eot");
}
LOD;
Main idea:
For more readability the pattern is divided into named subpatterns. The `(?(DEFINE)...)` doesn't match anything, it is only a definition section.
The main trick of this pattern is the use of the `\G` anchor that means: start of the string or contiguous to a precedent match. I added a negative lookbehind `(?<!^)` to avoid the first part of this definition.
The `<anchor>` named subpattern is the most important because it allows a match only if `@font-face {` is found or immediately after the end of an url (this is the reason why you can see a `["']?+`).
`<other_content>` represents all that is not an url section but matches url sections that must be skipped too(urls that begin with "http:", "data:"). The important detail of this subpattern is that it can't match the closing curly bracket of @font-face.
The mission of `<url_start>` is only to match `url("`.
`\K` resets all the substring that has been matched before from the match result.
`([^"'\s)}]*+)` matches the url (the only thing that stay in the match result with the leading `./../` )
Since `<other_content>` and the url subpattern can't match a `}` (that is outside quoted or comment parts), you are sure to never match something outside of the @font-face definition, the second consequence is that the pattern always fails after the last url. Thus, at the next attempt the "contiguous branch" will fail until the next @font-face.
another trick:
The main pattern begins with `\g<comment> (*SKIP)(*FAIL) |` to skip all content inside comments `/*....*/`. `\g<comment>` refers to the basic subpattern that describes how a comment look like. `(*SKIP)` forbids to retry the substring that has been matched before (on his left, by `g<comment>`), if the pattern fails on his right. `(*FAIL)` forces the pattern to fail. With this trick, comments are skipped and are not a match result (since the pattern fails).
subpatterns details:
quoted_content: It's used in `<other_content>` to avoid to match `url(` or `/*` that are inside quotes.
(["']) # capture group: the opening quote
(?> # atomic group: all possible content between quotes
[^"'\\]++ # all that is not a quote or a backslash
| # OR
\\{2} # two backslashes: (two \ doesn't escape anything)
| # OR
\\. # any escaped character
| # OR
(?!\g{-1})["'] # the other quote (this one that is not in the capture group)
)*+ # repeat zero or more time the atomic group
\g{-1} # backreference to the last capturing group
other_content: all that is not the closing curly bracket, or an url without `http:` or `data:`
(?> # open an atomic group
[^u}/"']++ # all character that are not problematic!
|
\g<quoted_content> # string inside quotes
|
\g<comment> # string inside comments
|
\Bu # "u" not preceded by a word boundary
|
u(?!rl\s*+\() # "u" not followed by "rl(" (not the start of an url definition)
|
/(?!\*) # "/" not followed by "*" (not the start of a comment)
|
\g<url_start> # match the url that begins with "http:"
\g<url_skip> ["']?+ # until the possible quote
)++ # repeat the atomic group one or more times
anchor
\G(?<!^) ["']?+ # contiguous to a precedent match with a possible closing quote
| # OR
@font-face \s*+ { # start of the @font-face definition
Notice:
You can improve the main pattern:
After the last url of @font-face, the regex engine attempts to match with the "contiguous branch" of `<anchor>` and match all characters until the `}` that makes the pattern fail. Then, on each same characters, the regex engine must try the two branches or `<anchor>` (that will always fail until the `}`.
To avoid these useless tries, you can change the main pattern to:
\g<comment> (*SKIP)(*FAIL) |
\g<anchor> \g<other_content>?+
(?>
\g<url_start> \K [./]*+ ([^"'\s)}]*+)
|
} (*SKIP)(*FAIL)
)
With this new scenario, the first character after the last url is matched by the "contiguous branch", `\g<other_content>` matches all until the `}`, `\g<url_start>` fails immediatly, the `}` is matched and `(*SKIP)(*FAIL)` make the pattern fail and forbids to retry these characters.
Problem
I have a web service that rewrites urls in css files so that they can be served via a CDN. The css files can contain urls to images or fonts. I currently have the following regex to match ALL urls within the css file: ``` (url\(\s*([\'\"]?+))((?!(https?\:|data\:|\.\.\/|\/))\S+)((\2)\s*\)) ``` However, I now want to introduce support for custom fonts and need to target the urls within `@font-fontface`: ``` @font-face { font-family: 'FontAwesome'; src: url("fonts/fontawesome-webfont.eot?v=4.0.3"); src: url("fonts/fontawesome-webfont.eot?#iefix&v=4.0.3") format("embedded-opentype"), url("fonts/fontawesome-webfont.woff?v=4.0.3") format("woff"), url("fonts/fontawesome-webfont.ttf?v=4.0.3") format("truetype"), url("fonts/fontawesome-webfont.svg?v=4.0.3#fontawesomeregular") format("svg"); font-weight: normal; font-style: normal; } ``` I then came up with the following: ``` @font-face\s*\{.*(url\(\s*([\'\"]?+))((?!(https?\:|data\:|\.\.\/|\/))\S+)((\2)\s*\))\s*\} ``` The problem is that this matches everything and not just the urls inside. I thought I can use lookbehind like so: ``` (?<=@font-face\s*\{.*)(url\(\s*([\'\"]?+))((?!(https?\:|data\:|\.\.\/|\/))\S+)((\2)\s*\))(?<=-\s*\}) ``` Unfortunately, PCRE (which PHP uses) does not support variable repetition within a lookbehind, so I am stuck. I do not wish to check for fonts by their extension as some fonts have the `.svg` extension which can conflict with images with the `.svg` extension. In addition, I would also like to modify my original regex to match all other urls that are NOT within an `@font-face`: ``` .someclass { background: url('images/someimage.png') no-repeat; } ``` Since I am unable to use lookbehinds, how can I extract the urls from those within a `@font-face` and those that are not within a `@font-face`?