Skip to main content

Python Regex Match 5 Digits Forex

Regular Expression HOWTO Dieses Dokument ist ein Einführungstutorial für die Verwendung von regulären Ausdrücken in Python mit dem re Modul. Es bietet eine sanftere Einführung als der entsprechende Abschnitt in der Bibliotheksreferenz. Einführung Das re-Modul wurde in Python 1.5 hinzugefügt und bietet Perl-artige reguläre Expressionsmuster. Frühere Versionen von Python kam mit dem regex-Modul, das Emacs-style Muster zur Verfügung stellte. Das Regex-Modul wurde komplett in Python 2.5 entfernt. Reguläre Ausdrücke (sogenannte REs oder Regexe oder Regex-Muster) sind im Wesentlichen eine winzige, hochspezialisierte Programmiersprache, die in Python eingebettet ist und über das re-Modul zur Verfügung gestellt wird. Wenn Sie diese kleine Sprache angeben, legen Sie die Regeln für die Menge der möglichen Zeichenfolgen fest, die Sie mit diesem Satz übereinstimmen möchten, können englische Sätze oder E-Mail-Adressen oder TeX-Befehle enthalten. Sie können dann Fragen wie 8220Does diese Saite entsprechen dem Muster8221, oder 8220Is dort eine Übereinstimmung für das Muster überall in diesem string8221. Sie können auch REs verwenden, um einen String zu modifizieren oder auf verschiedene Weise voneinander zu trennen. Reguläre Ausdrückmuster werden zu einer Reihe von Bytecodes kompiliert, die dann durch eine in C geschriebene übereinstimmende Engine ausgeführt werden. Für fortgeschrittene Anwendungen kann es notwendig sein, sorgfältig darauf zu achten, wie der Motor ein gegebenes RE ausführt und die RE in eine schreiben Bestimmte Weise, um einen schnelleren Bytecode zu erzeugen. Optimierung isn8217t in diesem Dokument abgedeckt, weil es erfordert, dass Sie ein gutes Verständnis der passenden Engine8217s Interna haben. Die Sprache des regulären Ausdrucks ist relativ klein und beschränkt, so dass nicht alle möglichen Zeichenfolgenverarbeitungsaufgaben mit regulären Ausdrücken durchgeführt werden können. Es gibt auch Aufgaben, die mit regulären Ausdrücken durchgeführt werden können, aber die Ausdrücke erweisen sich als sehr kompliziert. In diesen Fällen können Sie besser schreibe Python-Code, um die Verarbeitung zu tun, während Python-Code wird langsamer als ein aufwändiger regulärer Ausdruck, wird es wahrscheinlich auch mehr verständlich. Einfache Muster We8217ll beginnen, indem sie über die einfachsten möglichen regulären Ausdrücke lernen. Da reguläre Ausdrücke verwendet werden, um auf Strings zu operieren, beginnen wir mit der häufigsten Aufgabe: passende Zeichen. Für eine detaillierte Erklärung der Computerwissenschaft zugrunde liegenden regulären Ausdrücken (deterministische und nicht-deterministische endliche Automaten), können Sie sich auf fast jedem Lehrbuch auf Compiler schreiben. Passende Charaktere Die meisten Buchstaben und Buchstaben passen einfach zusammen. Beispielsweise wird der reguläre Ausdruckstest genau dem Stringtest entsprechen. (Sie können einen Groß - / Kleinschreibung-Modus aktivieren, der diesen RE-Test zu Test oder TEST sowie dazu später mehr ermöglichen würde.) Es gibt Ausnahmen von dieser Regel, dass einige Zeichen spezielle Metazeichen sind. Und don8217t passen sich. Stattdessen signalisieren sie, dass einige außergewöhnliche Dinge aufeinander abgestimmt sind oder sie andere Teile der RE betreffen, indem sie sie wiederholen oder ihre Bedeutung ändern. Ein Großteil dieses Dokuments widmet sich der Diskussion verschiedener Metazeichen und was sie tun. Hier wird eine vollständige Liste der Metazeichen ihrer Bedeutungen im restlichen HOWTO besprochen. Die ersten Metazeichen, die wir betrachten, sind und. Sie werden für die Angabe einer Zeichenklasse verwendet, bei der es sich um einen Satz von Zeichen handelt, die Sie anpassen möchten. Zeichen können einzeln aufgelistet werden, oder ein Zeichenbereich kann durch Angabe von zwei Zeichen und Trennung durch ein - angezeigt werden. Zum Beispiel passt abc zu einem der Zeichen a. B. Oder c ist das gleiche wie a-c. Die einen Bereich verwendet, um denselben Satz von Zeichen auszudrücken. Wenn Sie nur Kleinbuchstaben abgleichen möchten, wäre Ihr RE a-z. Metazeichen sind nicht in Klassen aktiv. Zum Beispiel wird akm mit einem der Zeichen a übereinstimmen. K. M Oder ist gewöhnlich ein metacharacter, aber innerhalb einer Charakterklasse it8217s, der von seiner speziellen Natur entkleidet ist. Sie können die Zeichen, die nicht in der Klasse aufgelistet sind, durch Ergänzung des Satzes übereinstimmen. Dies wird angezeigt, indem man a einschließt, da das erste Zeichen der Klasse außerhalb einer Zeichenklasse einfach mit dem Zeichen übereinstimmt. Zum Beispiel, 5 wird mit jedem Zeichen außer 5 übereinstimmen. Vielleicht ist das wichtigste Metazeichen der Backslash,. Wie in Python-String-Literalen können dem Backslash verschiedene Zeichen folgen, um verschiedene spezielle Sequenzen zu signalisieren. It8217s auch verwendet, um alle Metazeichen zu entgehen, so können Sie immer noch übereinstimmen sie in Mustern zum Beispiel, wenn Sie eine oder passen müssen. Können Sie ihnen einen Backslash voranstellen, um ihre spezielle Bedeutung zu entfernen: oder. Einige der speziellen Sequenzen beginnend mit repräsentieren vordefinierte Sätze von Zeichen, die oft nützlich sind, wie z. B. die Menge von Ziffern, die Menge der Buchstaben oder die Menge von allem, was isn8217t Whitespace. Die folgenden vordefinierten Spezialsequenzen sind eine Untermenge der verfügbaren. Die entsprechenden Klassen sind für Byte-String-Muster. Eine vollständige Liste der Sequenzen und erweiterten Klassendefinitionen für Unicode-Zeichenfolgenmuster finden Sie im letzten Teil der Regular Expression Syntax. D Entspricht einer Dezimalstelle, die der Klasse 0-9 entspricht. D Entspricht einem nicht stelligen Zeichen, das der Klasse 0-9 entspricht. S Entspricht allen Leerzeichen, die der Klasse tnrfv entsprechen. S Entspricht einem nicht-whitespace Zeichen, das der Klasse tnrfv entspricht. W Entspricht einem beliebigen alphanumerischen Zeichen, das der Klasse a-zA-Z0-9 entspricht. W Gleicht alle nicht-alphanumerischen Zeichen, die der Klasse a-zA-Z0-9 entsprechen. Diese Sequenzen können in einer Zeichenklasse enthalten sein. Zum Beispiel,. Ist eine Zeichenklasse, die mit jedem Leerraumzeichen übereinstimmt, oder, oder. . Die letzte Metazeichen in diesem Abschnitt ist. Es passt zu allem außer einem Zeilenumbruch-Zeichen, und there8217s ein alternativer Modus (re. DOTALL), wo es sogar eine Newline übereinstimmt. . Wird häufig verwendet, wo Sie 8220any character8221 entsprechen möchten. Wiederholen der Dinge In der Lage, unterschiedliche Mengen von Zeichen zu entsprechen, ist das erste, was reguläre Ausdrücke tun können, dass isn8217t bereits möglich mit den Methoden, die auf Strings. Wenn dies jedoch die einzige zusätzliche Fähigkeit von Registern war, würden sie keinen großen Fortschritt bedeuten. Eine andere Möglichkeit ist, dass Sie festlegen können, dass Teile der RE eine bestimmte Anzahl von Malen wiederholt werden müssen. Das erste Metazeichen für die Wiederholung von Dingen, die wir betrachten. Doesn8217t stimmt stattdessen mit dem Literalzeichen überein, es legt fest, dass das vorhergehende Zeichen null oder mehr mal anstatt genau einmal übereinstimmen kann. Zum Beispiel wird Katze ct (0 a Zeichen), cat (1 a), caaat (3 a Zeichen) und so weiter. Der RE-Motor hat verschiedene interne Beschränkungen, die sich aus der Größe von C8217s int-Typ, die es von übereinstimmenden über 2 Milliarden ein Zeichen, die Sie wahrscheinlich don8217t haben genug Speicher, um eine Zeichenfolge, die große, so dass Sie shouldn8217t in diese Grenze laufen zu verhindern. Wiederholungen wie gierig bei der Wiederholung eines RE, wird die passende Motor versuchen, es so oft wie möglich wiederholen. Wenn später Teile des Musters don8217t übereinstimmen, wird die passende Engine dann sichern und erneut versuchen, mit weniger Wiederholungen. Ein Schritt-für-Schritt-Beispiel wird dies deutlicher machen. Let8217s betrachten den Ausdruck abcdb. Dies entspricht dem Buchstaben a. Null oder mehr Buchstaben aus der Klasse bcd. Und endet schließlich mit einem b. Stellen Sie sich nun vor, diese RE gegen den String abcbd anzupassen. Versuchen Sie es erneut. Diesmal ist das Zeichen an der aktuellen Position b. So dass es gelingt. Das Ende der RE ist nun erreicht, und es hat abcb. Dies zeigt, wie die passende Engine geht so weit wie es kann, auf den ersten, und wenn keine Übereinstimmung gefunden wird, wird es dann schrittweise sichern und wiederholen Sie den Rest der RE wieder und wieder. Es wird sichern, bis es null Streichhölzer für bcd versucht hat. Und wenn das nachträglich fehlschlägt, wird der Motor schließen, dass der String doesn8217t mit dem RE überhaupt übereinstimmt. Ein anderes wiederholendes Metacharacter ist. Die ein oder mehrere Male übereinstimmt. Achten Sie sorgfältig auf die Differenz zwischen und passt null oder mehr mal, so dass, was immer wiederholt werden kann nicht vorhanden sein, während erfordert mindestens ein Vorkommen. Um ein ähnliches Beispiel zu verwenden, wird Katze Katze (1 a), caaat (3 a 8216s), aber won8217t Übereinstimmung ct. Es gibt zwei weitere Wiederholungsqualifikationen. Das Fragezeichenzeichen. Passt entweder einmal oder null mal können Sie es als Kennzeichnung etwas als optional denken. Zum Beispiel, home-brauen Spiele entweder homebrew oder home-brauen. Die komplizierteste wiederholte Qualifikation ist. Wobei m und n Dezimalzahlen sind. Dieser Qualifizierer bedeutet, dass es mindestens m Wiederholungen und höchstens n geben muss. Beispielsweise wird a / b mit a / b übereinstimmen. A / b. Und a / / b. Es gewann8217t Spiel ab. Die keinen Schrägstrich hat, oder ein //// b. Die vier hat. In diesem Fall können Sie entweder m oder n weglassen, für den fehlenden Wert wird ein vernünftiger Wert angenommen. Das Auslassen von m wird als untere Grenze von 0 interpretiert, während das Weglassen von n zu einer oberen Grenze der Unendlichkeit 8212 führt, tatsächlich ist die obere Grenze die 2-Milliardengrenze, die zuvor erwähnt wurde, aber das könnte auch unendlich sein. Die Leser eines Reduktionist gebogen kann bemerken, dass die drei anderen Qualifikationen können alle mit dieser Notation ausgedrückt werden. ist das gleiche wie . ist äquivalent zu . Und ist die gleiche wie. It8217s besser zu verwenden. . oder. Wenn Sie können, einfach weil they8217re kürzer und leichter zu lesen. Verwenden von regulären Ausdrücken Nun, da wir einige einfache reguläre Ausdrücke betrachtet haben, wie verwenden wir sie tatsächlich in Python. Das re Modul stellt eine Schnittstelle zu dem regulären Ausdruckmodul bereit, so dass Sie REs in Objekte kompilieren und dann Übereinstimmungen mit ihnen durchführen können. Kompilieren regulärer Ausdrücke Reguläre Ausdrücke werden in Musterobjekte kompiliert, die Methoden für verschiedene Operationen wie das Durchsuchen von Musterübereinstimmungen oder das Durchführen von Zeichenkettenersetzungen aufweisen. Repile () akzeptiert auch ein optionales Flags-Argument, um verschiedene Besonderheiten und Syntaxvariationen zu aktivieren. We8217ll gehen über die verfügbaren Einstellungen später, aber für jetzt ein einziges Beispiel zu tun: Die RE wird an repile () als String übergeben. REs werden als Strings behandelt, weil reguläre Ausdrücke aren8217t Teil der Kern-Python-Sprache sind und keine spezielle Syntax zum Ausdrücken erstellt wurde. (Es gibt Anwendungen, don8217t müssen REs überhaupt, so there8217s keine Notwendigkeit, aufblasen die Sprache Spezifikation, indem sie.) Stattdessen ist das re-Modul einfach ein C-Erweiterungsmodul mit Python, wie die Sockel oder zlib-Module enthalten. Putting REs in Strings hält die Python-Sprache einfacher, hat aber einen Nachteil, der das Thema des nächsten Abschnitts ist. Die Backslash-Pest Wie bereits erwähnt, verwenden reguläre Ausdrücke den umgekehrten Schrägstrich (), um spezielle Formulare anzugeben oder Sonderzeichen zu verwenden, ohne ihre spezielle Bedeutung aufzurufen. Diese Konflikte mit Python8217s Verwendung des gleichen Zeichens für den gleichen Zweck in Zeichenfolgenliteralen. Let8217s sagen, dass Sie eine RE schreiben möchten, die mit dem String-Abschnitt übereinstimmt. Die in einer LaTeX-Datei gefunden werden können. Um herauszufinden, was in dem Programmcode zu schreiben, mit der gewünschten Zeichenfolge abgestimmt werden. Als Nächstes müssen Sie Backslashs und andere Metazeichen entfernen, indem Sie ihnen einen Backslash voranstellen, der zum String-Abschnitt führt. Der resultierende String, der an repile () übergeben werden muss, muss sein. Um dies jedoch als Python-String-Literal auszudrücken, müssen beide Backslashs erneut entschlüsselt werden. Finden Sie alle Teilstrings, in denen die RE übereinstimmt, und gibt sie als Iterator zurück. Match () und search () return Keine, wenn keine Übereinstimmung gefunden werden kann. Wenn they8217re erfolgreich ist, wird eine Match-Objekt-Instanz zurückgegeben, die Informationen über die Übereinstimmung enthält: wo es beginnt und endet, dessen Teilstring und mehr. Sie können dies durch interaktives Experimentieren mit dem re Modul erfahren. Wenn Sie Tkinter verfügbar haben, können Sie auch Tools / scripts / redemo. py ansehen. Ein Demonstrationsprogramm, das in der Python-Distribution enthalten ist. Es erlaubt Ihnen, REs und Zeichenfolgen einzugeben, und zeigt an, ob das RE übereinstimmt oder ausfällt. Redemo. py kann sehr nützlich sein, wenn man versucht, eine komplizierte RE zu debuggen. Phil Schwartz8217s Kodos ist auch ein interaktives Werkzeug zur Entwicklung und Erprobung von RE-Mustern. Dieses HOWTO verwendet den Standard-Python-Interpreter für seine Beispiele. Führen Sie zunächst den Python-Interpreter aus, importieren Sie das re-Modul und kompilieren Sie ein RE: Jetzt können Sie versuchen, verschiedene Zeichenfolgen mit dem RE a-z zu vergleichen. Eine leere Zeichenfolge shouldn8217t übereinstimmen, da bedeutet 8216one oder mehr Wiederholungen8217. Match () sollte in diesem Fall keine zurückgeben, was dazu führt, dass der Interpreter keine Ausgabe druckt. Sie können explizit das Ergebnis von match () ausgeben, um dies klar zu machen. Nun, let8217s versuchen es auf eine Zeichenfolge, die es, wie Tempo übereinstimmen sollte. In diesem Fall gibt match () ein Match-Objekt zurück. So dass Sie das Ergebnis in einer Variablen für spätere Verwendung speichern sollte. Nun können Sie das Match-Objekt nach Informationen über den passenden String abfragen. Match-Objekt-Instanzen haben auch mehrere Methoden und Attribute die wichtigsten sind: Der Versuch dieser Methoden wird bald klären, ihre Bedeutung: group () gibt die Teilzeichenfolge, die von der RE übereinstimmen. Start () und end () geben den Anfangs - und Endindex der Übereinstimmung zurück. Span () gibt sowohl Start - als auch End-Indizes in einem einzigen Tupel zurück. Da die match () - Methode nur überprüft, ob die RE zu Beginn einer Zeichenfolge übereinstimmt, wird start () immer null. Allerdings durchsucht die search () - Methode von Mustern den String, so dass die Übereinstimmung in diesem Fall nicht mit Null beginnen kann. In den eigentlichen Programmen ist die häufigste Art, das Match-Objekt in einer Variablen zu speichern, und dann überprüfen, ob es None war. Dies sieht normalerweise wie folgt aus: Zwei Mustermethoden geben alle Übereinstimmungen für ein Muster zurück. Findall () gibt eine Liste von übereinstimmenden Zeichenfolgen zurück: findall () muss die gesamte Liste erstellen, bevor sie als Ergebnis zurückgegeben werden kann. Die finditer () - Methode gibt eine Sequenz von Match-Objektinstanzen als Iterator zurück. 1 Module-Level-Funktionen Sie müssen ein Musterobjekt erstellen und seine Methoden aufrufen. Das re-Modul bietet auch Top-Level-Funktionen namens match (). Suche(). finde alle(). Sub (). und so weiter. Diese Funktionen verwenden dieselben Argumente wie die entsprechende Mustermethode, wobei die RE-Zeichenfolge als erstes Argument hinzugefügt wurde und noch keine oder keine Objektobjekt-Instanz zurückgegeben wird. Unter der Haube, diese Funktionen erstellen Sie einfach ein Muster-Objekt für Sie und rufen Sie die entsprechende Methode auf sie. Sie speichern auch das kompilierte Objekt in einem Cache, so dass zukünftige Anrufe mit dem gleichen RE sind schneller. Verwenden Sie diese Module-Ebene Funktionen, oder sollten Sie das Muster und rufen Sie seine Methoden selbst Diese Wahl hängt davon ab, wie oft die RE verwendet werden, und auf Ihre persönliche Codierung Stil. Wenn das RE an nur einem Punkt im Code verwendet wird, dann sind die Modulfunktionen wahrscheinlich bequemer. Wenn ein Programm viele reguläre Ausdrücke enthält oder dieselben an mehreren Stellen erneut verwendet, dann könnte es sinnvoll sein, alle Definitionen an einem Ort zu sammeln, in einem Abschnitt des Codes, der alle REs vor der Zeit kompiliert. Um ein Beispiel aus der Standardbibliothek zu nehmen, gibt es hier einen Auszug aus dem veralteten xmllib-Modul: Ich ziehe es vor, mit dem kompilierten Objekt zu arbeiten, auch für einmalige Anwendungen, aber nur wenige Menschen werden so viel puristisch sein, wie ich bin . Kompilierungsflaggen Zusammenstellungsflags können Sie einige Aspekte von regulären Ausdrücken ändern. Flags sind im re-Modul unter zwei Namen, einem langen Namen wie IGNORECASE und einem kurzen, ein-Buchstaben-Formular wie I. (Wenn you8217 vertraut mit Perl8217s Muster-Modifikatoren sind, verwenden die Ein-Buchstaben-Formulare die gleichen Buchstaben die kurze Form Von re. VERBOSE ist zB re. X.) Mehrere Flags können durch bitweises ODER-Verknüpfen spezifiziert werden. I re. M setzt z. B. die I - und M-Flags. Hier ist eine Tabelle der verfügbaren Flaggen, gefolgt von einer ausführlicheren Erklärung von jedem. Macht mehrere Escapes wie w. B. S und d abhängig von der Unicode-Zeichen-Datenbank. Führen Sie case-insensitive passende Zeichenklasse und literale Zeichenfolgen übereinstimmen Buchstaben von Ignorieren Fall. Zum Beispiel wird A-Z Kleinbuchstaben entsprechen, und Spam wird Spam entsprechen. Spam. Oder Spam. Diese Lowercasing doesn8217t nehmen die aktuelle Gebietsschema berücksichtigt werden, wenn Sie auch die LOCALE-Flag gesetzt. Machen w. W. b. Und B. abhängig vom aktuellen Gebietsschema. Locales sind ein Merkmal der C-Bibliothek, die dazu beitragen soll, Programme zu schreiben, die Sprachunterschiede berücksichtigen. Zum Beispiel, wenn you8217re Verarbeitung Französisch Text, you8217d wollen in der Lage, w schreiben, um Wörter zu entsprechen, aber w entspricht nur der Zeichenklasse A-Za-z es gewonnen8217t übereinstimmen oder. Wenn Ihr System richtig konfiguriert ist und ein französisches Gebietsschema ausgewählt ist, zeigen bestimmte C-Funktionen das Programm an, das auch als Buchstabe betrachtet werden sollte. Das Setzen des LOCALE-Flags beim Kompilieren eines regulären Ausdrucks bewirkt, dass das resultierende kompilierte Objekt diese C-Funktionen für w verwendet, das langsamer ist, aber auch ermöglicht, w französische Wörter zu vergleichen, wie you8217d erwarten. (Und haven8217t wurde erklärt, aber they8217ll werden im Abschnitt Weitere Metazeichen eingeführt.) Meistens passt nur am Anfang des Strings und passt nur am Ende des Strings und unmittelbar vor dem Zeilenende (falls vorhanden) am Ende des Strings. Wenn dieses Flag angegeben ist, passt es am Anfang des Strings und am Anfang jeder Zeile innerhalb des Strings, unmittelbar nach jedem Newline. Ebenso passt das Metazeichen entweder am Ende des Strings und am Ende jeder Zeile (unmittelbar vor jedem Zeilenumbruch). Macht das . Sonderzeichen entsprechen einem beliebigen Zeichen, einschließlich eines Zeilenumbruchs ohne dieses Flag,. Wird alles außer einem Zeilenumbruch passen. Machen w. W. b. B. d. D. s und S abhängig von der Unicode-Zeichen-Eigenschaften-Datenbank. Mit dieser Markierung können Sie reguläre Ausdrücke schreiben, die besser lesbar sind, indem Sie mehr Flexibilität bei der Formatierung erhalten. Wenn dieses Flag angegeben wurde, wird der Whitespace innerhalb der RE-Zeichenfolge ignoriert, es sei denn, der Whitespace befindet sich in einer Zeichenklasse oder wird von einem unverschobenen Backslash vorangestellt. Dieses Flag können Sie auch Kommentare innerhalb einer RE, die von der Engine ignoriert werden Kommentare werden durch ein that8217s weder in einer Zeichenklasse oder durch einen unverschobenen Backslash vorausgeht markiert. Zum Beispiel hier8217s eine RE, die re. VERBOSE verwendet, um zu sehen, wie viel einfacher es zu lesen ist Ohne die ausführliche Einstellung würde die RE wie folgt aussehen: Im obigen Beispiel wurde Python8217s automatische Verkettung von String-Literalen verwendet, um die RE zu brechen In kleinere Stücke, aber es8217s noch schwieriger zu verstehen als die Version mit re. VERBOSE. Mehr Pattern Power Bisher haben wir nur einen Teil der Features von regulären Ausdrücken abgedeckt. In diesem Abschnitt behandeln wir einige neue Metazeichen und wie man Gruppen verwendet, um Teile des Textes abzurufen, der abgestimmt wurde. Mehr Metazeichen Es gibt einige Metazeichen, die wir noch nicht behandelt haben. Die meisten von ihnen werden in diesem Abschnitt behandelt werden. Einige der verbleibenden Metazeichen, die diskutiert werden sollen, sind nullbreitige Aussagen. Sie don8217t Ursache der Motor, um durch die Zeichenfolge statt, sie verbrauchen keine Zeichen überhaupt, und einfach Erfolg oder Misserfolg. Beispielsweise ist b eine Aussage, dass sich die aktuelle Position an einer Wortgrenze befindet, wobei die Position n8217t durch das b überhaupt geändert wird. Das bedeutet, dass nullinterne Aussagen nie wiederholt werden sollten, da sie, wenn sie einmal an einem gegebenen Ort übereinstimmen, offensichtlich unendlich viele Male angeglichen werden können. Alternation oder der 8220or8221-Operator. Wenn A und B reguläre Ausdrücke sind, entspricht AB jedem beliebigen String, der entweder A oder B entspricht, eine sehr geringe Priorität, um die Arbeit sinnvoller zu machen, wenn Sie mehrere Zeichenfolgen abwechseln. CrowServo entspricht entweder Crow oder Servo. Nicht Cro. Ein w oder ein S. und ervo. Um ein Literal. benutzen . Oder in eine Zeichenklasse einschließen, wie in. Übereinstimmungen am Zeilenanfang. Wenn das Flag MULTILINE nicht gesetzt ist, wird es nur am Anfang des Strings übereinstimmen. Im MULTILINE-Modus entspricht dies auch unmittelbar nach jeder Newline innerhalb der Zeichenfolge. Wenn Sie beispielsweise das Wort From nur am Anfang einer Zeile anpassen möchten, ist das zu verwendende RE von From. Entspricht dem Ende einer Zeile, die als das Ende des Strings definiert ist, oder an beliebiger Stelle, gefolgt von einem Zeilenumbruchzeichen. Um ein Literal. Verwenden oder in eine Zeichenklasse einschließen, wie in. A Stimmt nur am Anfang des Strings. Wenn nicht im MULTILINE-Modus, A und sind wirksam die gleichen. Im MULTILINE-Modus sind sie unterschiedlich: A stimmt immer noch nur am Anfang des Strings überein, kann aber an jeder beliebigen Stelle innerhalb der Zeichenfolge übereinstimmen, die einem Zeilenumbruchzeichen folgt. Z Gilt nur am Ende des Strings. B Wortgrenze. Dies ist eine Behauptung von null Breiten, die nur am Anfang oder Ende eines Wortes passt. Ein Wort wird als eine Folge von alphanumerischen Zeichen definiert, so dass das Ende eines Wortes durch Leerzeichen oder ein nicht-alphanumerisches Zeichen angezeigt wird. Das folgende Beispiel stimmt mit der Klasse nur überein, wenn it8217s ein vollständiges Wort, won8217t übereinstimmt, wenn it8217s in einem anderen Wort enthalten ist. Es gibt zwei Feinheiten, die Sie beachten sollten, wenn Sie diese spezielle Sequenz verwenden. Erstens, dies ist die schlimmste Kollision zwischen Python8217s String-Literalen und regulären Ausdruck Sequenzen. In Python8217s String-Literalen ist b das Backspace-Zeichen, ASCII-Wert 8. Wenn you8217re keine rohen Strings verwendet, konvertiert Python das b in einen Backspace und Ihr RE won8217t übereinstimmen, wie Sie es erwarten. Das folgende Beispiel sieht genauso aus wie unser vorheriges RE, aber weglässt das r vor der RE-Zeichenfolge. Zweitens, innerhalb einer Zeichenklasse, wo there8217s keine Verwendung für diese Behauptung, b repräsentiert das Backspace-Zeichen, für die Kompatibilität mit Python8217s String-Literale. B Eine andere Null-Breite-Aussage, dies ist das Gegenteil von b. Wenn die aktuelle Position nicht an einer Wortgrenze liegt. Gruppierung Häufig müssen Sie mehr Informationen erhalten als nur, ob die RE übereinstimmt oder nicht. Reguläre Ausdrücke werden oft verwendet, um Zeichenfolgen zu sezieren, indem sie ein RE schreiben, das in mehrere Untergruppen unterteilt ist, die mit verschiedenen interessierenden Komponenten übereinstimmen. Beispielsweise wird eine RFC-822-Kopfzeile in einen Header-Namen und einen Wert aufgeteilt, getrennt durch:. Wie folgt: Dies kann gehandhabt werden, indem ein regulärer Ausdruck erstellt wird, der einer ganzen Kopfzeile entspricht und eine Gruppe hat, die mit dem Header-Namen übereinstimmt, und eine andere Gruppe, die dem header8217s-Wert entspricht. Gruppen werden durch die (.) - Metazeichen markiert. (Und) die gleiche Bedeutung haben wie in mathematischen Ausdrücken, gruppieren sie die in ihnen enthaltenen Ausdrücke und können den Inhalt einer Gruppe mit einem sich wiederholenden Qualifikationsmerkmal wie z. B. wiederholen. . oder . Zum Beispiel wird (ab) mit null oder mehr Wiederholungen von ab übereinstimmen. Gruppen, die mit (.) Gekennzeichnet sind, erfassen auch den Anfangs - und Endindex des Textes, mit dem sie übereinstimmen, indem sie ein Argument an group () übergeben. Anfang(). Ende(). Und span (). Gruppen sind nummeriert, beginnend mit 0. Gruppe 0 ist immer vorhanden it8217s die ganze RE, so passen Objekt-Methoden alle haben Gruppe 0 als ihr Standard-Argument. Später we8217ll sehen, wie Gruppen auszudrücken, dass don8217t erfassen die Spanne von Text, dass sie übereinstimmen. Untergruppen sind von links nach rechts, von 1 nach oben nummeriert. Gruppen können verschachtelt werden, um die Zahl zu bestimmen, zählen Sie einfach die öffnenden Klammerzeichen, von links nach rechts. Group () können mehrere Gruppennummern gleichzeitig übergeben werden. In diesem Fall wird ein Tupel zurückgegeben, das die entsprechenden Werte für diese Gruppen enthält. Die group () - Methode gibt ein Tupel zurück, das die Zeichenketten für alle Untergruppen enthält, von 1 bis zu vielen, die es gibt. Rückreferenzen in einem Muster können Sie festlegen, dass der Inhalt einer früheren Erfassungsgruppe auch an der aktuellen Position im String gefunden werden muss. Beispielsweise wird 1 erfolgreich sein, wenn der genaue Inhalt der Gruppe 1 an der aktuellen Position gefunden werden kann und ansonsten fehlschlägt. Beachten Sie, dass Python8217s String-Literale auch einen Backslash verwenden, gefolgt von Zahlen, um beliebige Zeichen in einen String zu integrieren. Daher sollten Sie unbedingt einen Raw-String verwenden, wenn Sie Rückreferenzen in einem RE integrieren. Beispielsweise erkennt das folgende RE doppelte Wörter in einem String. Backreferences wie diese aren8217t oft nützlich für nur die Suche durch eine Zeichenfolge 8212 gibt es wenige Textformate, die Daten auf diese Weise 8212 wiederholen, aber you8217ll bald herausfinden, dass they8217re sehr nützlich bei der Durchführung von String-Substitutionen. Non-Capture und benannte Gruppen Aufwendige REs können viele Gruppen verwenden, sowohl um Teilstrings von Interesse zu erfassen, als auch um die RE selbst zu gruppieren und zu strukturieren. In komplexen REs wird es schwierig, die Gruppe Zahlen zu verfolgen. Es gibt zwei Funktionen, die mit diesem Problem helfen. Beide verwenden eine allgemeine Syntax für reguläre Ausdruckserweiterungen, also sehen wir uns das zuerst an. Perl 5 hat zusätzliche reguläre Ausdrücke hinzugefügt, und das Python re Modul unterstützt die meisten von ihnen. Es wäre schwierig gewesen, neue Single-Keystroke-Metazeichen oder neue spezielle Sequenzen auszuwählen, die mit der Darstellung der neuen Features beginnen, ohne Perl8217s reguläre Ausdrücke verwirrend von Standard-REs zu unterscheiden. Wenn Sie Amp als ein neues Metazeichen wählen, würden zum Beispiel alte Ausdrücke davon ausgehen, dass der Verstärker ein regelmäßiger Charakter war und es nicht durch Schreiben von amp oder amp entkommen konnte. Die von den Perl-Entwicklern gewählte Lösung war die Verwendung von (.) Als Erweiterungssyntax. Unmittelbar nachdem eine Klammer ein Syntaxfehler war, weil die. Hätte nichts zu wiederholen, so dass dieses didn8217t keine Kompatibilitätsprobleme einzuführen. Die Zeichen unmittelbar nach der. Geben Sie an, welche Erweiterung verwendet wird, also ist (foo) eine Sache (eine positive lookahead-Behauptung) und (: foo) etwas anderes (eine nicht erfassende Gruppe, die den Unterausdruck foo enthält). Python fügt eine Erweiterungssyntax zur Perl8217s-Erweiterungssyntax hinzu. Wenn das erste Zeichen nach dem Fragezeichen ein P. Sie wissen, dass es8217s eine Erweiterung that8217s spezifisch für Python. Derzeit gibt es zwei solche Erweiterungen: (Pltnamegt.) Definiert eine benannte Gruppe, und (Pname) ist eine Rückbeziehung zu einer benannten Gruppe. Wenn zukünftige Versionen von Perl 5 ähnliche Funktionen mit einer anderen Syntax hinzufügen, wird das re-Modul geändert, um die neue Syntax zu unterstützen, während die python-spezifische Syntax für compatibility8217s sake bewahrt wird. Nachdem wir nun die allgemeine Erweiterungssyntax betrachtet haben, können wir auf die Funktionen zurückgreifen, die die Arbeit mit Gruppen in komplexen REs vereinfachen. Da Gruppen von links nach rechts nummeriert werden und ein komplexer Ausdruck viele Gruppen verwenden kann, kann es schwierig werden, die korrekte Nummerierung zu verfolgen. Das Ändern einer solchen komplexen RE ist auch ärgerlich: Fügen Sie eine neue Gruppe in der Nähe des Anfangs ein, und Sie ändern die Nummern von allem, was darauf folgt. Manchmal möchten Sie eine Gruppe verwenden, um einen Teil eines regulären Ausdrucks zu sammeln, aber aren8217t daran interessiert, den Inhalt der group8217s abzurufen. Sie können diese Tatsache explizit mit einer nicht erfassenden Gruppe machen: (.). Wo Sie die ersetzen können. Mit jedem anderen regulären Ausdruck. Abgesehen von der Tatsache, dass Sie den Inhalt von dem, was die Gruppe zusammengestellt hat, abrufen können, verhält sich eine Nicht-Capturing-Gruppe genauso wie eine Capturing-Gruppe, die Sie alles in ihr setzen können. Wiederholen Sie sie mit einem Wiederholungs-Metazeichen. Und nisten es in anderen Gruppen (Capturing oder Non-Capture). (.) Ist besonders nützlich, wenn Sie ein bestehendes Muster ändern, da Sie neue Gruppen hinzufügen können, ohne zu ändern, wie alle anderen Gruppen numeriert sind. Es sollte erwähnt werden, dass es keinen Leistungsunterschied bei der Suche zwischen Capturing und Non-Capture-Gruppen gibt, weder Form noch schneller als die andere. Ein wichtigeres Merkmal heißt Gruppen: Anstatt sich auf sie durch Zahlen zu beziehen, können Gruppen durch einen Namen referenziert werden. Die Syntax für eine benannte Gruppe ist eine der Python-spezifischen Erweiterungen: (Pltnamegt.). Name ist natürlich der Name der Gruppe. Benannte Gruppen verhalten sich genauso wie Capturing-Gruppen und verknüpfen zusätzlich einen Namen mit einer Gruppe. Die Match-Objektmethoden, die mit Capture-Gruppen umgehen, akzeptieren entweder ganze Zahlen, die sich auf die Gruppe nach Anzahl oder Zeichenfolgen beziehen, die den gewünschten group8217s-Namen enthalten. Named Gruppen sind immer noch Zahlen, so können Sie Informationen über eine Gruppe auf zwei Arten abrufen: Named Gruppen sind praktisch, weil sie Sie leicht-erinnern Namen verwenden können, anstatt Zahlen zu merken. Hier8217s ein Beispiel RE aus dem imaplib-Modul: It8217s offensichtlich viel einfacher, m. group (zonem) abzurufen. Anstatt sich daran zu erinnern, die Gruppe 9 abzurufen. Die Syntax für Rückreferenzen in einem Ausdruck wie (.) 1 bezieht sich auf die Nummer der Gruppe. There8217s natürlich eine Variante, die den Gruppennamen anstelle der Zahl verwendet. Dies ist eine weitere Python-Erweiterung: (Pname), die angibt, dass der Inhalt der Gruppe namens name wieder am aktuellen Punkt abgeglichen werden soll. Der reguläre Ausdruck für die Suche nach verdoppelten Wörtern (bw) s1 kann auch als (Plwordgtbw) s (Pword) geschrieben werden: Lookahead-Assertionen Eine weitere Nulldurchsetzungs-Assertion ist die Lookahead-Assertion. Lookahead-Behauptungen sind sowohl in positiver als auch in negativer Form verfügbar und sehen folgendermaßen aus: (.) ​​Positive Voraussicht. Dies gelingt, wenn der enthaltene reguläre Ausdruck, hier dargestellt durch. Erfolgreich an der aktuellen Position übereinstimmt und ansonsten fehlschlägt. Aber, sobald der enthaltene Ausdruck ausprobiert worden ist, wird die passende Engine nicht allmählich den Rest des Musters voranbringen, wo die Assertion gestartet wird. (.) Negative Voraussicht. Dies ist das Gegenteil der positiven Assertion, die es gelingt, wenn der enthaltene Ausdruck doesn8217t an der aktuellen Position in der Zeichenfolge übereinstimmt. Um diese konkrete, let8217s Blick auf einen Fall, wo ein lookahead nützlich ist. Betrachten Sie ein einfaches Muster, das einem Dateinamen entspricht, und teilen Sie es auseinander in einen Basisnamen und eine Erweiterung, getrennt durch ein. Zum Beispiel in news. rc. News ist der Basisname, und rc ist die Dateiname8217s-Erweiterung. Das dazu passende Muster ist ganz einfach: Beachten Sie, dass die. Muss speziell behandelt werden, weil es ein metacharacter I8217ve legte es in einer Charakterklasse. Beachten Sie auch, dass das Hinzufügen hinzugefügt wird, um sicherzustellen, dass der gesamte String in der Erweiterung enthalten sein muss. Dieser reguläre Ausdruck stimmt mit foo. bar und autoexec. bat und sendmail. cf und printers. conf überein. Nun, betrachten kompliziert das Problem ein wenig, was, wenn Sie Dateinamen, in denen die Erweiterung ist nicht bat Match. Einige fehlerhafte Versuche:.b. Der erste Versuch oben versucht, bat auszuschließen, indem es erfordert, dass das erste Zeichen der Erweiterung nicht b ist. Das ist falsch, weil das Muster auch tutn8217t mit foo. bar übereinstimmt. Der Ausdruck wird unübersichtlicher, wenn Sie versuchen, die erste Lösung aufzurüsten, indem Sie einen der folgenden Fälle benötigen: das erste Zeichen der Erweiterung isn8217t b das zweite Zeichen isn8217t a oder das dritte Zeichen isn8217t t. Das akzeptiert foo. bar und verwirft autoexec. bat. Aber es erfordert eine Drei-Buchstaben-Erweiterung und won8217t akzeptieren einen Dateinamen mit einer Zwei-Buchstaben-Erweiterung wie sendmail. cf. We8217ll komplizieren das Muster wieder in dem Bemühen, es zu beheben. In the third attempt, the second and third letters are all made optional in order to allow matching extensions shorter than three characters, such as sendmail. cf . The pattern8217s getting really complicated now, which makes it hard to read and understand. Worse, if the problem changes and you want to exclude both bat and exe as extensions, the pattern would get even more complicated and confusing. A negative lookahead cuts through all this confusion: .(bat). The negative lookahead means: if the expression bat doesn8217t match at this point, try the rest of the pattern if bat does match, the whole pattern will fail. The trailing is required to ensure that something like sample. batch. where the extension only starts with bat. will be allowed. The . makes sure that the pattern works when there are multiple dots in the filename. Excluding another filename extension is now easy simply add it as an alternative inside the assertion. The following pattern excludes filenames that end in either bat or exe : Modifying Strings Up to this point, we8217ve simply performed searches against a static string. Regular expressions are also commonly used to modify strings in various ways, using the following pattern methods: Splitting Strings The split() method of a pattern splits a string apart wherever the RE matches, returning a list of the pieces. It8217s similar to the split() method of strings but provides much more generality in the delimiters that you can split by split() only supports splitting by whitespace or by a fixed string. As you8217d expect, there8217s a module-level re. split() function, too. Split string by the matches of the regular expression. If capturing parentheses are used in the RE, then their contents will also be returned as part of the resulting list. If maxsplit is nonzero, at most maxsplit splits are performed. You can limit the number of splits made, by passing a value for maxsplit . When maxsplit is nonzero, at most maxsplit splits will be made, and the remainder of the string is returned as the final element of the list. In the following example, the delimiter is any sequence of non-alphanumeric characters. Sometimes you8217re not only interested in what the text between delimiters is, but also need to know what the delimiter was. If capturing parentheses are used in the RE, then their values are also returned as part of the list. Compare the following calls: The module-level function re. split() adds the RE to be used as the first argument, but is otherwise the same. Search and Replace Another common task is to find all the matches for a pattern, and replace them with a different string. The sub() method takes a replacement value, which can be either a string or a function, and the string to be processed. Returns the string obtained by replacing the leftmost non-overlapping occurrences of the RE in string by the replacement replacement . If the pattern isn8217t found, string is returned unchanged. The optional argument count is the maximum number of pattern occurrences to be replaced count must be a non-negative integer. The default value of 0 means to replace all occurrences. Here8217s a simple example of using the sub() method. It replaces colour names with the word colour : The subn() method does the same work, but returns a 2-tuple containing the new string value and the number of replacements that were performed: Empty matches are replaced only when they8217re not adjacent to a previous match. If replacement is a string, any backslash escapes in it are processed. That is, n is converted to a single newline character, r is converted to a carriage return, and so forth. Unknown escapes such as j are left alone. Backreferences, such as 6. are replaced with the substring matched by the corresponding group in the RE. This lets you incorporate portions of the original text in the resulting replacement string. This example matches the word section followed by a string enclosed in . and changes section to subsection : There8217s also a syntax for referring to named groups as defined by the (Pltnamegt. ) syntax. gltnamegt will use the substring matched by the group named name. and gltnumbergt uses the corresponding group number. glt2gt is therefore equivalent to 2. but isn8217t ambiguous in a replacement string such as glt2gt0. ( 20 would be interpreted as a reference to group 20, not a reference to group 2 followed by the literal character 0 .) The following substitutions are all equivalent, but use all three variations of the replacement string. replacement can also be a function, which gives you even more control. If replacement is a function, the function is called for every non-overlapping occurrence of pattern . On each call, the function is passed a match object argument for the match and can use this information to compute the desired replacement string and return it. In the following example, the replacement function translates decimals into hexadecimal: When using the module-level re. sub() function, the pattern is passed as the first argument. The pattern may be provided as an object or as a string if you need to specify regular expression flags, you must either use a pattern object as the first parameter, or use embedded modifiers in the pattern string, e. g. sub(quot(i)bquot, quotxquot, quotbbbb BBBBquot) returns x x . Common Problems Regular expressions are a powerful tool for some applications, but in some ways their behaviour isn8217t intuitive and at times they don8217t behave the way you may expect them to. This section will point out some of the most common pitfalls. Use String Methods Sometimes using the re module is a mistake. If you8217re matching a fixed string, or a single character class, and you8217re not using any re features such as the IGNORECASE flag, then the full power of regular expressions may not be required. Strings have several methods for performing operations with fixed strings and they8217re usually much faster, because the implementation is a single small C loop that8217s been optimized for the purpose, instead of the large, more generalized regular expression engine. One example might be replacing a single fixed string with another one for example, you might replace word with deed. re. sub() seems like the function to use for this, but consider the replace() method. Note that replace() will also replace word inside words, turning swordfish into sdeedfish. but the naive RE word would have done that, too. (To avoid performing the substitution on parts of words, the pattern would have to be bwordb. in order to require that word have a word boundary on either side. This takes the job beyond replace() 8216s abilities.) Another common task is deleting every occurrence of a single character from a string or replacing it with another single character. You might do this with something like re. sub(n, , S). but translate() is capable of doing both tasks and will be faster than any regular expression operation can be. In short, before turning to the re module, consider whether your problem can be solved with a faster and simpler string method. match() versus search() The match() function only checks if the RE matches at the beginning of the string while search() will scan forward through the string for a match. It8217s important to keep this distinction in mind. Remember, match() will only report a successful match which will start at 0 if the match wouldn8217t start at zero, match() will not report it. On the other hand, search() will scan forward through the string, reporting the first match it finds. Sometimes you8217ll be tempted to keep using re. match(). and just add . to the front of your RE. Resist this temptation and use re. search() instead. The regular expression compiler does some analysis of REs in order to speed up the process of looking for a match. One such analysis figures out what the first character of a match must be for example, a pattern starting with Crow must match starting with a C. The analysis lets the engine quickly scan through the string looking for the starting character, only trying the full match if a C is found. Adding . defeats this optimization, requiring scanning to the end of the string and then backtracking to find a match for the rest of the RE. Use re. search() instead. Greedy versus Non-Greedy When repeating a regular expression, as in a. the resulting action is to consume as much of the pattern as possible. This fact often bites you when you8217re trying to match a pair of balanced delimiters, such as the angle brackets surrounding an HTML tag. The naive pattern for matching a single HTML tag doesn8217t work because of the greedy nature of . . The RE matches the lt in lthtmlgt. and the . consumes the rest of the string. There8217s still more left in the RE, though, and the gt can8217t match at the end of the string, so the regular expression engine has to backtrack character by character until it finds a match for the gt. The final match extends from the lt in lthtmlgt to the gt in lt/titlegt. which isn8217t what you want. In this case, the solution is to use the non-greedy qualifiers . . or . which match as little text as possible. In the above example, the gt is tried immediately after the first lt matches, and when it fails, the engine advances a character at a time, retrying the gt at every step. This produces just the right result: (Note that parsing HTML or XML with regular expressions is painful. Quick-and-dirty patterns will handle common cases, but HTML and XML have special cases that will break the obvious regular expression by the time you8217ve written a regular expression that handles all of the possible cases, the patterns will be very complicated. Use an HTML or XML parser module for such tasks.) Using re. VERBOSE By now you8217ve probably noticed that regular expressions are a very compact notation, but they8217re not terribly readable. REs of moderate complexity can become lengthy collections of backslashes, parentheses, and metacharacters, making them difficult to read and understand. For such REs, specifying the re. VERBOSE flag when compiling the regular expression can be helpful, because it allows you to format the regular expression more clearly. The re. VERBOSE flag has several effects. Whitespace in the regular expression that isn8217t inside a character class is ignored. This means that an expression such as dog cat is equivalent to the less readable dogcat. but a b will still match the characters a. B. or a space. In addition, you can also put comments inside a RE comments extend from a character to the next newline. When used with triple-quoted strings, this enables REs to be formatted more neatly: This is far more readable than: Feedback Regular expressions are a complicated topic. Did this document help you understand them Were there parts that were unclear, or Problems you encountered that weren8217t covered here If so, please send suggestions for improvements to the author. The most complete book on regular expressions is almost certainly Jeffrey Friedl8217s Mastering Regular Expressions, published by O8217Reilly. Unfortunately, it exclusively concentrates on Perl and Java8217s flavours of regular expressions, and doesn8217t contain any Python material at all, so it won8217t be useful as a reference for programming in Python. (The first edition covered Python8217s now-removed regex module, which won8217t help you much.) Consider checking it out from your library. Introduced in Python 2.2.2.7.2. re 8212 Regular expression operations This module provides regular expression matching operations similar to those found in Perl. Both patterns and strings to be searched can be Unicode strings as well as 8-bit strings. However, Unicode strings and 8-bit strings cannot be mixed: that is, you cannot match an Unicode string with a byte pattern or vice-versa similarly, when asking for a substitution, the replacement string must be of the same type as both the pattern and the search string. Regular expressions use the backslash character ( ) to indicate special forms or to allow special characters to be used without invoking their special meaning. This collides with Python8217s usage of the same character for the same purpose in string literals for example, to match a literal backslash, one might have to write as the pattern string, because the regular expression must be . and each backslash must be expressed as inside a regular Python string literal. The solution is to use Python8217s raw string notation for regular expression patterns backslashes are not handled in any special way in a string literal prefixed with r . So rquotnquot is a two-character string containing and n . while quotnquot is a one-character string containing a newline. Usually patterns will be expressed in Python code using this raw string notation. It is important to note that most regular expression operations are available as module-level functions and RegexObject methods. The functions are shortcuts that don8217t require you to compile a regex object first, but miss some fine-tuning parameters. Mastering Regular Expressions Book on regular expressions by Jeffrey Friedl, published by O8217Reilly. The second edition of the book no longer covers Python at all, but the first edition covered writing good regular expression patterns in great detail. 7.2.1. Regular Expression Syntax A regular expression (or RE) specifies a set of strings that matches it the functions in this module let you check if a particular string matches a given regular expression (or if a given regular expression matches a particular string, which comes down to the same thing). Regular expressions can be concatenated to form new regular expressions if A and B are both regular expressions, then AB is also a regular expression. In general, if a string p matches A and another string q matches B . the string pq will match AB. This holds unless A or B contain low precedence operations boundary conditions between A and B or have numbered group references. Thus, complex expressions can easily be constructed from simpler primitive expressions like the ones described here. For details of the theory and implementation of regular expressions, consult the Friedl book referenced above, or almost any textbook about compiler construction. A brief explanation of the format of regular expressions follows. For further information and a gentler presentation, consult the Regular Expression HOWTO . Regular expressions can contain both special and ordinary characters. Most ordinary characters, like A . a . or 0 . are the simplest regular expressions they simply match themselves. You can concatenate ordinary characters, so last matches the string last . (In the rest of this section, we8217ll write RE8217s in this special style . usually without quotes, and strings to be matched in single quotes .) Some characters, like or ( . are special. Special characters either stand for classes of ordinary characters, or affect how the regular expressions around them are interpreted. Regular expression pattern strings may not contain null bytes, but can specify the null byte using the number notation, e. g. x00 . The special characters are: . (Dot.) In the default mode, this matches any character except a newline. If the DOTALL flag has been specified, this matches any character including a newline. (Caret.) Matches the start of the string, and in MULTILINE mode also matches immediately after each newline. Matches the end of the string or just before the newline at the end of the string, and in MULTILINE mode also matches before a newline. foo matches both 8216foo8217 and 8216foobar8217, while the regular expression foo matches only 8216foo8217. More interestingly, searching for foo. in foo1nfoo2n matches 8216foo28217 normally, but 8216foo18217 in MULTILINE mode searching for a single in foon will find two (empty) matches: one just before the newline, and one at the end of the string. Causes the resulting RE to match 0 or more repetitions of the preceding RE, as many repetitions as are possible. ab will match 8216a8217, 8216ab8217, or 8216a8217 followed by any number of 8216b8217s. Causes the resulting RE to match 1 or more repetitions of the preceding RE. ab will match 8216a8217 followed by any non-zero number of 8216b8217s it will not match just 8216a8217. Causes the resulting RE to match 0 or 1 repetitions of the preceding RE. ab will match either 8216a8217 or 8216ab8217. . . . The . . and qualifiers are all greedy they match as much text as possible. Sometimes this behaviour isn8217t desired if the RE lt. gt is matched against ltH1gttitlelt/H1gt . it will match the entire string, and not just ltH1gt . Adding after the qualifier makes it perform the match in non-greedy or minimal fashion as few characters as possible will be matched. Using . in the previous expression will match only ltH1gt . Specifies that exactly m copies of the previous RE should be matched fewer matches cause the entire RE not to match. For example, a will match exactly six a characters, but not five. Causes the resulting RE to match from m to n repetitions of the preceding RE, attempting to match as many repetitions as possible. For example, a will match from 3 to 5 a characters. Omitting m specifies a lower bound of zero, and omitting n specifies an infinite upper bound. As an example, a b will match aaaab or a thousand a characters followed by a b . but not aaab . The comma may not be omitted or the modifier would be confused with the previously described form. Causes the resulting RE to match from m to n repetitions of the preceding RE, attempting to match as few repetitions as possible. This is the non-greedy version of the previous qualifier. For example, on the 6-character string aaaaaa . a will match 5 a characters, while a will only match 3 characters. Either escapes special characters (permitting you to match characters like . . and so forth), or signals a special sequence special sequences are discussed below. If you8217re not using a raw string to express the pattern, remember that Python also uses the backslash as an escape sequence in string literals if the escape sequence isn8217t recognized by Python8217s parser, the backslash and subsequent character are included in the resulting string. However, if Python would recognize the resulting sequence, the backslash should be repeated twice. This is complicated and hard to understand, so it8217s highly recommended that you use raw strings for all but the simplest expressions. Used to indicate a set of characters. Characters can be listed individually, or a range of characters can be indicated by giving two characters and separating them by a - . Special characters are not active inside sets. For example, akm will match any of the characters a . k . m . or a-z will match any lowercase letter, and a-zA-Z0-9 matches any letter or digit. Character classes such as w or S (defined below) are also acceptable inside a range, although the characters they match depends on whether ASCII or LOCALE mode is in force. If you want to include a or a - inside a set, precede it with a backslash, or place it as the first character. The pattern will match . beispielsweise. You can match the characters not within a range by complementing the set. This is indicated by including a as the first character of the set elsewhere will simply match the character. For example, 5 will match any character except 5 . and will match any character except . Note that inside the special forms and special characters lose their meanings and only the syntaxes described here are valid. For example, . . ( . ) . and so on are treated as literals inside . and backreferences cannot be used inside . AB . where A and B can be arbitrary REs, creates a regular expression that will match either A or B. An arbitrary number of REs can be separated by the in this way. This can be used inside groups (see below) as well. As the target string is scanned, REs separated by are tried from left to right. When one pattern completely matches, that branch is accepted. This means that once A matches, B will not be tested further, even if it would produce a longer overall match. In other words, the operator is never greedy. To match a literal . use . or enclose it inside a character class, as in . (. ) Matches whatever regular expression is inside the parentheses, and indicates the start and end of a group the contents of a group can be retrieved after a match has been performed, and can be matched later in the string with the number special sequence, described below. To match the literals ( or ) . use ( or ) . or enclose them inside a character class: ( ) . (. ) This is an extension notation (a following a ( is not meaningful otherwise). The first character after the determines what the meaning and further syntax of the construct is. Extensions usually do not create a new group (Pltnamegt. ) is the only exception to this rule. Following are the currently supported extensions. (aiLmsux) (One or more letters from the set a . i . L . m . s . u . x .) The group matches the empty string the letters set the corresponding flags: re. A (ASCII-only matching), re. I (ignore case), re. L (locale dependent), re. M (multi-line), re. S (dot matches all), and re. X (verbose), for the entire regular expression. (The flags are described in Module Contents .) This is useful if you wish to include the flags as part of the regular expression, instead of passing a flag argument to the repile() function. Note that the (x) flag changes how the expression is parsed. It should be used first in the expression string, or after one or more whitespace characters. If there are non-whitespace characters before the flag, the results are undefined. (. ) A non-capturing version of regular parentheses. Matches whatever regular expression is inside the parentheses, but the substring matched by the group cannot be retrieved after performing a match or referenced later in the pattern. (Pltnamegt. ) Similar to regular parentheses, but the substring matched by the group is accessible within the rest of the regular expression via the symbolic group name name . Group names must be valid Python identifiers, and each group name must be defined only once within a regular expression. A symbolic group is also a numbered group, just as if the group were not named. So the group named id in the example below can also be referenced as the numbered group 1 . For example, if the pattern is (Pltidgta-zA-Zw) . the group can be referenced by its name in arguments to methods of match objects, such as m. group(id) or m. end(id) . and also by name in the regular expression itself (using (Pid) ) and replacement text given to. sub() (using gltidgt ). (Pname) Matches whatever text was matched by the earlier group named name . (. ) A comment the contents of the parentheses are simply ignored. (. ) Matches if . matches next, but doesn8217t consume any of the string. This is called a lookahead assertion. For example, Isaac (Asimov) will match Isaac only if it8217s followed by Asimov . (. ) Matches if . doesn8217t match next. This is a negative lookahead assertion. For example, Isaac (Asimov) will match Isaac only if it8217s not followed by Asimov . (lt. ) Matches if the current position in the string is preceded by a match for . that ends at the current position. This is called a positive lookbehind assertion . (ltabc)def will find a match in abcdef . since the lookbehind will back up 3 characters and check if the contained pattern matches. The contained pattern must only match strings of some fixed length, meaning that abc or ab are allowed, but a and a are not. Note that patterns which start with positive lookbehind assertions will never match at the beginning of the string being searched you will most likely want to use the search() function rather than the match() function: This example looks for a word following a hyphen: (lt. ) Matches if the current position in the string is not preceded by a match for . . This is called a negative lookbehind assertion . Similar to positive lookbehind assertions, the contained pattern must only match strings of some fixed length. Patterns which start with negative lookbehind assertions may match at the beginning of the string being searched. ((id/name)yes-patternno-pattern) Will try to match with yes-pattern if the group with given id or name exists, and with no-pattern if it doesn8217t. no-pattern is optional and can be omitted. For example, (lt)(w64w(:.w))((1)gt) is a poor email matching pattern, which will match with ltuser64hostgt as well as user64host . but not with ltuser64host nor user64hostgt . The special sequences consist of and a character from the list below. If the ordinary character is not on the list, then the resulting RE will match the second character. For example, matches the character . number Matches the contents of the group of the same number. Groups are numbered starting from 1. For example, (.) 1 matches the the or 55 55 . but not the end (note the space after the group). This special sequence can only be used to match one of the first 99 groups. If the first digit of number is 0, or number is 3 octal digits long, it will not be interpreted as a group match, but as the character with octal value number . Inside the and of a character class, all numeric escapes are treated as characters. A Matches only at the start of the string. b Matches the empty string, but only at the beginning or end of a word. A word is defined as a sequence of Unicode alphanumeric or underscore characters, so the end of a word is indicated by whitespace or a non-alphanumeric, non-underscore Unicode character. Note that formally, b is defined as the boundary between a w and a W character (or vice versa). By default Unicode alphanumerics are the ones used, but this can be changed by using the ASCII flag. Inside a character range, b represents the backspace character, for compatibility with Python8217s string literals. B Matches the empty string, but only when it is not at the beginning or end of a word. This is just the opposite of b . so word characters are Unicode alphanumerics or the underscore, although this can be changed by using the ASCII flag. d For Unicode (str) patterns: Matches any Unicode digit (which includes 0-9 . and also many other digit characters). If the ASCII flag is used only 0-9 is matched (but the flag affects the entire regular expression, so in such cases using an explicit 0-9 may be a better choice). For 8-bit (bytes) patterns: Matches any decimal digit this is equivalent to 0-9 . D Matches any character which is not a Unicode decimal digit. This is the opposite of d . If the ASCII flag is used this becomes the equivalent of 0-9 (but the flag affects the entire regular expression, so in such cases using an explicit 0-9 may be a better choice). s For Unicode (str) patterns: Matches Unicode whitespace characters (which includes tnrfv . and also many other characters, for example the non-breaking spaces mandated by typography rules in many languages). If the ASCII flag is used, only tnrfv is matched (but the flag affects the entire regular expression, so in such cases using an explicit tnrfv may be a better choice). For 8-bit (bytes) patterns: Matches characters considered whitespace in the ASCII character set this is equivalent to tnrfv . S Matches any character which is not a Unicode whitespace character. This is the opposite of s . If the ASCII flag is used this becomes the equivalent of tnrfv (but the flag affects the entire regular expression, so in such cases using an explicit tnrfv may be a better choice). w For Unicode (str) patterns: Matches Unicode word characters this includes most characters that can be part of a word in any language, as well as numbers and the underscore. If the ASCII flag is used, only a-zA-Z0-9 is matched (but the flag affects the entire regular expression, so in such cases using an explicit a-zA-Z0-9 may be a better choice). For 8-bit (bytes) patterns: Matches characters considered alphanumeric in the ASCII character set this is equivalent to a-zA-Z0-9 . W Matches any character which is not a Unicode word character. This is the opposite of w . If the ASCII flag is used this becomes the equivalent of a-zA-Z0-9 (but the flag affects the entire regular expression, so in such cases using an explicit a-zA-Z0-9 may be a better choice). Z Matches only at the end of the string. Most of the standard escapes supported by Python string literals are also accepted by the regular expression parser: Octal escapes are included in a limited form: If the first digit is a 0, or if there are three octal digits, it is considered an octal escape. Otherwise, it is a group reference. As for string literals, octal escapes are always at most three digits in length. 7.2.2. Matching vs Searching Python offers two different primitive operations based on regular expressions: match checks for a match only at the beginning of the string, while search checks for a match anywhere in the string (this is what Perl does by default). Note that match may differ from search even when using a regular expression beginning with . matches only at the start of the string, or in MULTILINE mode also immediately following a newline. The 8220match8221 operation succeeds only if the pattern matches at the start of the string regardless of mode, or at the starting position given by the optional pos argument regardless of whether a newline precedes it. 7.2.3. Module Contents The module defines several functions, constants, and an exception. Some of the functions are simplified versions of the full featured methods for compiled regular expressions. Most non-trivial applications always use the compiled form. Compile a regular expression pattern into a regular expression object, which can be used for matching using its match() and search() methods, described below. The expression8217s behaviour can be modified by specifying a flags value. Values can be any of the following variables, combined using bitwise OR (the operator). is equivalent to but using repile() and saving the resulting regular expression object for reuse is more efficient when the expression will be used several times in a single program. The compiled versions of the most recent patterns passed to re. match() . re. search() or repile() are cached, so programs that use only a few regular expressions at a time needn8217t worry about compiling regular expressions. Make w . W . b . B . d . D . s and S perform ASCII-only matching instead of full Unicode matching. This is only meaningful for Unicode patterns, and is ignored for byte patterns. Note that for backward compatibility, the re. U flag still exists (as well as its synonym re. UNICODE and its embedded counterpart (u) ), but these are redundant in Python 3 since matches are Unicode by default for strings (and Unicode matching isn8217t allowed for bytes). re. I re. IGNORECASE Perform case-insensitive matching expressions like A-Z will match lowercase letters, too. This is not affected by the current locale and works for Unicode characters as expected. re. L re. LOCALE Make w . W . b . B . s and S dependent on the current locale. The use of this flag is discouraged as the locale mechanism is very unreliable, and it only handles one 8220culture8221 at a time anyway you should use Unicode matching instead, which is the default in Python 3 for Unicode (str) patterns. re. M re. MULTILINE When specified, the pattern character matches at the beginning of the string and at the beginning of each line (immediately following each newline) and the pattern character matches at the end of the string and at the end of each line (immediately preceding each newline). By default, matches only at the beginning of the string, and only at the end of the string and immediately before the newline (if any) at the end of the string. re. S re. DOTALL Make the . special character match any character at all, including a newline without this flag, . will match anything except a newline. re. X re. VERBOSE This flag allows you to write regular expressions that look nicer. Whitespace within the pattern is ignored, except when in a character class or preceded by an unescaped backslash, and, when a line contains a neither in a character class or preceded by an unescaped backslash, all characters from the leftmost such through the end of the line are ignored. That means that the two following regular expression objects that match a decimal number are functionally equal: re. search ( pattern . string . flags ) Scan through string looking for a location where the regular expression pattern produces a match, and return a corresponding MatchObject instance. Return None if no position in the string matches the pattern note that this is different from finding a zero-length match at some point in the string. re. match ( pattern . string . flags ) If zero or more characters at the beginning of string match the regular expression pattern . return a corresponding MatchObject instance. Return None if the string does not match the pattern note that this is different from a zero-length match. If you want to locate a match anywhere in string . use search() instead. Split string by the occurrences of pattern . If capturing parentheses are used in pattern . then the text of all groups in the pattern are also returned as part of the resulting list. If maxsplit is nonzero, at most maxsplit splits occur, and the remainder of the string is returned as the final element of the list. If there are capturing groups in the separator and it matches at the start of the string, the result will start with an empty string. The same holds for the end of the string: That way, separator components are always found at the same relative indices within the result list (e. g. if there8217s one capturing group in the separator, the 0th, the 2nd and so forth). Note that split will never split a string on an empty pattern match. For example: Changed in version 3.1: Added the optional flags argument. re. findall ( pattern . string . flags ) Return all non-overlapping matches of pattern in string . as a list of strings. The string is scanned left-to-right, and matches are returned in the order found. If one or more groups are present in the pattern, return a list of groups this will be a list of tuples if the pattern has more than one group. Empty matches are included in the result unless they touch the beginning of another match. re. finditer ( pattern . string . flags ) Return an iterator yielding MatchObject instances over all non-overlapping matches for the RE pattern in string . The string is scanned left-to-right, and matches are returned in the order found. Empty matches are included in the result unless they touch the beginning of another match. re. sub ( pattern . repl . string . count . flags ) Return the string obtained by replacing the leftmost non-overlapping occurrences of pattern in string by the replacement repl . If the pattern isn8217t found, string is returned unchanged. repl can be a string or a function if it is a string, any backslash escapes in it are processed. That is, n is converted to a single newline character, r is converted to a linefeed, and so forth. Unknown escapes such as j are left alone. Backreferences, such as 6 . are replaced with the substring matched by group 6 in the pattern. For example: If repl is a function, it is called for every non-overlapping occurrence of pattern . The function takes a single match object argument, and returns the replacement string. For example: The pattern may be a string or an RE object. The optional argument count is the maximum number of pattern occurrences to be replaced count must be a non-negative integer. If omitted or zero, all occurrences will be replaced. Empty matches for the pattern are replaced only when not adjacent to a previous match, so sub(x, -, abc) returns - a-b-c - . In addition to character escapes and backreferences as described above, gltnamegt will use the substring matched by the group named name . as defined by the (Pltnamegt. ) syntax. gltnumbergt uses the corresponding group number glt2gt is therefore equivalent to 2 . but isn8217t ambiguous in a replacement such as glt2gt0 . 20 would be interpreted as a reference to group 20, not a reference to group 2 followed by the literal character 0 . The backreference glt0gt substitutes in the entire substring matched by the RE. Changed in version 3.1: Added the optional flags argument. Perform the same operation as sub() . but return a tuple (newstring, numberofsubsmade) . Changed in version 3.1: Added the optional flags argument. re. escape ( string ) Return string with all non-alphanumerics backslashed this is useful if you want to match an arbitrary literal string that may have regular expression metacharacters in it. re. purge ( ) Clear the regular expression cache. exception re. error Exception raised when a string passed to one of the functions here is not a valid regular expression (for example, it might contain unmatched parentheses) or when some other error occurs during compilation or matching. It is never an error if a string contains no match for a pattern. 7.2.4. Regular Expression Objects The RegexObject class supports the following methods and attributes: Scan through string looking for a location where this regular expression produces a match, and return a corresponding MatchObject instance. Return None if no position in the string matches the pattern note that this is different from finding a zero-length match at some point in the string. The optional second parameter pos gives an index in the string where the search is to start it defaults to 0 . This is not completely equivalent to slicing the string the pattern character matches at the real beginning of the string and at positions just after a newline, but not necessarily at the index where the search is to start. The optional parameter endpos limits how far the string will be searched it will be as if the string is endpos characters long, so only the characters from pos to endpos - 1 will be searched for a match. If endpos is less than pos . no match will be found, otherwise, if rx is a compiled regular expression object, rx. search(string, 0, 50) is equivalent to rx. search(string:50, 0) . If zero or more characters at the beginning of string match this regular expression, return a corresponding MatchObject instance. Return None if the string does not match the pattern note that this is different from a zero-length match. The optional pos and endpos parameters have the same meaning as for the search() method. If you want to locate a match anywhere in string . use search() instead. split ( string . maxsplit0 ) Identical to the split() function, using the compiled pattern. findall ( string . pos . endpos ) Similar to the findall() function, using the compiled pattern, but also accepts optional pos and endpos parameters that limit the search region like for match() . finditer ( string . pos . endpos ) Similar to the finditer() function, using the compiled pattern, but also accepts optional pos and endpos parameters that limit the search region like for match() . sub ( repl . string . count0 ) Identical to the sub() function, using the compiled pattern. subn ( repl . string . count0 ) Identical to the subn() function, using the compiled pattern. flags The flags argument used when the RE object was compiled, or 0 if no flags were provided. groups The number of capturing groups in the pattern. groupindex A dictionary mapping any symbolic group names defined by (Pltidgt) to group numbers. The dictionary is empty if no symbolic groups were used in the pattern. pattern The pattern string from which the RE object was compiled. 7.2.5. Match Objects Match Objects always have a boolean value of True . so that you can test whether e. g. match() resulted in a match with a simple if statement. They support the following methods and attributes: expand ( template ) Return the string obtained by doing backslash substitution on the template string template . as done by the sub() method. Escapes such as n are converted to the appropriate characters, and numeric backreferences ( 1 . 2 ) and named backreferences ( glt1gt . gltnamegt ) are replaced by the contents of the corresponding group. group ( group1 . . ) Returns one or more subgroups of the match. If there is a single argument, the result is a single string if there are multiple arguments, the result is a tuple with one item per argument. Without arguments, group1 defaults to zero (the whole match is returned). If a groupN argument is zero, the corresponding return value is the entire matching string if it is in the inclusive range 1..99, it is the string matching the corresponding parenthesized group. If a group number is negative or larger than the number of groups defined in the pattern, an IndexError exception is raised. If a group is contained in a part of the pattern that did not match, the corresponding result is None . If a group is contained in a part of the pattern that matched multiple times, the last match is returned. If the regular expression uses the (Pltnamegt. ) syntax, the groupN arguments may also be strings identifying groups by their group name. If a string argument is not used as a group name in the pattern, an IndexError exception is raised. A moderately complicated example: 7.2.6.3. Avoiding recursion If you create regular expressions that require the engine to perform a lot of recursion, you may encounter a RuntimeError exception with the message maximum recursion limit exceeded. For example, You can often restructure your regular expression to avoid recursion. Simple uses of the pattern are special-cased to avoid recursion. Thus, the above regular expression can avoid recursion by being recast as Begin a-zA-Z0-9 end . As a further benefit, such regular expressions will run faster than their recursive equivalents. 7.2.6.4. search() vs. match() In a nutshell, match() only attempts to match a pattern at the beginning of a string where search() will match a pattern anywhere in a string. For example: The following applies only to regular expression objects like those created with repile(quotpatternquot) . not the primitives re. match(pattern, string) or re. search(pattern, string) . match() has an optional second parameter that gives an index in the string where the search is to start: 7.2.6.5. Making a Phonebook split() splits a string into a list delimited by the passed pattern. The method is invaluable for converting textual data into data structures that can be easily read and modified by Python as demonstrated in the following example that creates a phonebook. First, here is the input. Normally it may come from a file, here we are using triple-quoted string syntax: The entries are separated by one or more newlines. Now we convert the string into a list with each nonempty line having its own entry: Finally, split each entry into a list with first name, last name, telephone number, and address. We use the maxsplit parameter of split() because the address has spaces, our splitting pattern, in it: The . pattern matches the colon after the last name, so that it does not occur in the result list. With a maxsplit of 4 . we could separate the house number from the street name: 7.2.6.6. Text Munging sub() replaces every occurrence of a pattern with a string or the result of a function. This example demonstrates using sub() with a function to 8220munge8221 text, or randomize the order of all the characters in each word of a sentence except for the first and last characters: 7.2.6.7. Finding all Adverbs findall() matches all occurrences of a pattern, not just the first one as search() does. For example, if one was a writer and wanted to find all of the adverbs in some text, he or she might use findall() in the following manner: 7.2.6.8. Finding all Adverbs and their Positions If one wants more information about all matches of a pattern than the matched text, finditer() is useful as it provides instances of MatchObject instead of strings. Continuing with the previous example, if one was a writer who wanted to find all of the adverbs and their positions in some text, he or she would use finditer() in the following manner: 7.2.6.9. Raw String Notation Raw string notation ( rquottextquot ) keeps regular expressions sane. Without it, every backslash ( ) in a regular expression would have to be prefixed with another one to escape it. For example, the two following lines of code are functionally identical: When one wants to match a literal backslash, it must be escaped in the regular expression. With raw string notation, this means rquotquot . Without raw string notation, one must use quotquot . making the following lines of code functionally identical:Python Regular Expressions A regular expression is a special sequence of characters that helps you match or find other strings or sets of strings, using a specialized syntax held in a pattern. Regular expressions are widely used in UNIX world. The module re provides full support for Perl-like regular expressions in Python. The re module raises the exception re. error if an error occurs while compiling or using a regular expression. We would cover two important functions, which would be used to handle regular expressions. But a small thing first: There are various characters, which would have special meaning when they are used in regular expression. To avoid any confusion while dealing with regular expressions, we would use Raw Strings as rexpression . The match Function This function attempts to match RE pattern to string with optional flags . Here is the syntax for this function minus Here is the description of the parameters: This is the regular expression to be matched. This is the string, which would be searched to match the pattern at the beginning of string. You can specify different flags using bitwise OR (). These are modifiers, which are listed in the table below. The re. match function returns a match object on success, None on failure. We use group(num) or groups() function of match object to get matched expression. Match Object Methods This method returns entire match (or specific subgroup num) This method returns all matching subgroups in a tuple (empty if there werent any) Example When the above code is executed, it produces following result minus The search Function This function searches for first occurrence of RE pattern within string with optional flags . Here is the syntax for this function: Here is the description of the parameters: This is the regular expression to be matched. This is the string, which would be searched to match the pattern anywhere in the string. You can specify different flags using bitwise OR (). These are modifiers, which are listed in the table below. The re. search function returns a match object on success, none on failure. We use group(num) or groups() function of match object to get matched expression. Match Object Methods This method returns entire match (or specific subgroup num) This method returns all matching subgroups in a tuple (empty if there werent any) Example When the above code is executed, it produces following result minus Matching Versus Searching Python offers two different primitive operations based on regular expressions: match checks for a match only at the beginning of the string, while search checks for a match anywhere in the string (this is what Perl does by default). Example When the above code is executed, it produces the following result minus Search and Replace One of the most important re methods that use regular expressions is sub . Syntax This method replaces all occurrences of the RE pattern in string with repl . substituting all occurrences unless max provided. This method returns modified string. Example When the above code is executed, it produces the following result minus Regular Expression Modifiers: Option Flags Regular expression literals may include an optional modifier to control various aspects of matching. The modifiers are specified as an optional flag. You can provide multiple modifiers using exclusive OR (), as shown previously and may be represented by one of these minus


Comments

Popular posts from this blog

Bollinger Bands Expert Advisor Mt4

MetaTrader 4 - Indikatoren Bollinger Bands, BB - Indikator für MetaTrader 4 Beschreibung: Bollinger Bands Technische Indikator (BB) ähnelt Umschlägen. Der einzige Unterschied besteht darin, dass die Bänder von Umschlägen einen festen Abstand () von dem gleitenden Durchschnitt entfernt sind, während die Bollinger-Bänder eine bestimmte Anzahl von Standardabweichungen von ihr weg aufgetragen sind. Die Standardabweichung ist ein Maß für die Volatilität, daher passen sich die Bollinger-Bänder den Marktbedingungen an. Wenn die Märkte volatiler werden, erweitern sich die Bands und schließen sich in weniger volatilen Perioden zusammen. Bollinger Bands sind in der Regel auf der Preisliste aufgetragen, können aber auch dem Indikatordiagramm hinzugefügt werden (Custom Indicators). Wie bei den Envelopes basiert die Interpretation der Bollinger-Bands darauf, dass die Preise zwischen der oberen und unteren Zeile der Bands bleiben. Eine Besonderheit des Bollinger Band Indikators ist seine variable Br...

Financemalta Devisenmarkt

Ein Devisen-Service ist genehmigungs nach dem Investment Services Act (im Folgenden als ISA bezeichnet), Kapitel 370 der Gesetze von Malta und ist passportable Rahmen der EU-Märkte für Finanzinstrumente (im Folgenden als MiFID genannt), wenn der Dienst bezieht sich auf Verträge über die Differenz, Derivate in Bezug auf Devisen und Rolling Spot Forex. einen Service in Bezug auf Bereitstellung von Devisen, die zu Anlagezwecken erworben oder gehalten wird, ist unter dem ISA auch genehmigungs jedoch ist dies eine MiFID-Service nicht berücksichtigt und daher nicht im Rahmen dieser EU-Richtlinie notifiziert werden kann. Der Service des Online-Devisenhandels wird in der Regel in einer von zwei Formen angeboten, und zwar entweder durch eigene Geschäfte oder als risikoloser Principal, oft als White Label Partner. Für Letztere würde das Unternehmen mit einem anderen Haupt bei der Ausführung von zwei passenden Trades, ein mit dem Kunden und ein Gegengeschäft beteiligt sein, trat zur gleichen Zeit...

Punkt Und Abbildung Forex Pdf Strategie

Grundlegendes zu Point-Figure-Charts Teil I von IV - Die Geschichte von Point-Amp-Figure-Charts - Warum Trader Verwendung von Point-Amp-Figuren-Charts - Der Aufbau von Point-Amp-Figur-Charts Die Point-Amp-Figur - oder PampF-Charts sind einzigartig für jede andere, Dieser Artikel wird Ihnen bequem mit dem Charting-Technik, die über 100 Jahre alt ist. Ähnlich den japanischen Candlesticks haben die Point-Amp-Figuren-Charts den Test der Zeit gestanden, da sie die Erkennung von Staus oder Trendausbrüchen einfach machen. Die Geschichte der Point-Amp-Figuren-Charts Im Gegensatz zu anderen Charting-Methoden gibt es keine Person, die für die Erstellung von Point-Amp-Figuren-Charts gutgeschrieben wird. Vor Computern wurden Punkt-Amp-Figuren von einer Methode, die von Boden-Trader in der 19. und Pre-Computer 20. Jahrhundert. Die grundlegende Prämisse, dass PampF geboren wurde, ist, dass es notwendig, eine einfache Methode für Boden Händler, um Preis-Aktion zu analysieren Preis ohne unnötigen Lärm...